Frequency Analysis
Statistical analysis of character frequency in ciphertext, useful for cracking classical substitution ciphers.
Input Text
Instructions
Analysis Principle
Frequency analysis is a classic cryptanalysis method based on the statistical regularity of character occurrence in natural languages. For example, in English the letter E has the highest frequency (about 12.7%), followed by T, A, O, I, N, S, H, R. By comparing the frequency distribution of ciphertext and plaintext, one can infer substitution cipher mappings.
Statistics Items
Total characters and unique characters reflect text size and complexity; information entropy measures the uncertainty of character distribution, higher entropy means more random; character type distribution categorizes by uppercase, lowercase, digits, whitespace, punctuation, and other; bigram frequency shows co-occurrence patterns of adjacent character pairs, useful for identifying language features.
Steps
1. Paste text to analyze in the input box; 2. System automatically calculates character frequency and displays statistics overview; 3. View character type distribution and frequency ranking; 4. Use tabs above frequency ranking to toggle between 'By Frequency' and 'By Character' sorting; 5. View bigram frequency to assist cryptanalysis.
Notes
Shorter text leads to larger frequency distribution deviation and lower reference value; recommended to input at least 100 characters. Maximum entropy is log₂(alphabet size); pure random text entropy approaches maximum. This tool is suitable for auxiliary analysis of classical ciphers like substitution and transposition ciphers, not effective against modern encryption algorithms.