Letters
Letter frequency in English
The A–Z table, where the standard figures come from, why no two sources agree exactly, and what the numbers are actually used for.
Letter frequency is one of those figures that everybody has seen and few people have seen a source for. The numbers below come from analyses of large corpora of English prose; they are reliable to within a few tenths of a per cent, and the caveats afterwards explain why greater precision than that is not really available.
The A–Z table
| Letter | Frequency | Rank | Letter | Frequency | Rank |
|---|---|---|---|---|---|
| E | 12.70% | 1 | M | 2.41% | 14 |
| T | 9.06% | 2 | W | 2.36% | 15 |
| A | 8.17% | 3 | F | 2.23% | 16 |
| O | 7.51% | 4 | G | 2.02% | 17 |
| I | 6.97% | 5 | Y | 1.97% | 18 |
| N | 6.75% | 6 | P | 1.93% | 19 |
| S | 6.33% | 7 | B | 1.49% | 20 |
| H | 6.09% | 8 | V | 0.98% | 21 |
| R | 5.99% | 9 | K | 0.77% | 22 |
| D | 4.25% | 10 | J | 0.15% | 23 |
| L | 4.03% | 11 | X | 0.15% | 24 |
| C | 2.78% | 12 | Q | 0.10% | 25 |
| U | 2.76% | 13 | Z | 0.07% | 26 |
The famous mnemonic for the top twelve is ETAOIN SHRDLU, which was not invented as a mnemonic at all. It is the order of the first two columns of a Linotype machine's keyboard, arranged by frequency so that the most common letters were nearest the operator's hands. Typesetters who slipped would fill a line with it, and it turns up as a printing error in a surprising number of old newspapers.
Vowels and consonants
Adding the five vowels gives 38.11 per cent, which is where the familiar “about 40 per cent vowels” figure comes from. Add Y and it rises to just over 40. The remaining letters make up the other 60 per cent, and five consonants — T, N, S, H and R — account for more than half of that on their own.
This is why a Scrabble rack with one vowel is a bad rack and why constructing a crossword grid is mostly a problem of vowel placement.
Position matters as much as frequency
Where a letter appears changes its usefulness enormously. The most common first letters of English words are, in order: T, A, O, I, S. E, which dominates overall, is only the tenth most common word-initial letter, because it is overwhelmingly an interior and terminal letter.
The most common final letters are E, S, D, T, N — E from silent terminal e, S from plurals and third-person verbs, D and T from past tenses.
Why no two tables agree exactly
The figures above are typical of general English prose. They shift according to what was counted:
- Dictionary counts versus text counts. Counting each headword once gives a very different table from counting running text, where the alone contributes about seven per cent of all words.
- Genre. Legal writing pushes up C, S and T. Chemistry pushes up Y and X. Text messages push up U and R for reasons that have nothing to do with English.
- Period. Nineteenth-century prose has a measurably different distribution from contemporary web writing.
- Proper nouns. A corpus of news is full of names, which distorts J, K and Z upwards.
Treat any table as accurate to a few tenths of a per cent and no further.
What people use this for
- Breaking substitution ciphers. Frequency analysis is the classical first move, described by al-Kindi in the ninth century and still the foundation of every introductory cryptanalysis course.
- Word games. Scrabble tile distribution and point values are derived directly from letter frequency — Z and Q are worth ten points because they are rare. Wordle strategy is the same problem restricted to five-letter words.
- Keyboard layout. Dvorak and Colemak are both attempts to place frequent letters on the home row; QWERTY famously is not.
- Typography and type design. The relative frequency of letters determines which letterforms deserve the most attention in a new typeface.
- Teaching phonics. High-frequency letters and combinations are taught first for good reason.
Check your own text
The A–Z comb below every counter on this site shows the distribution in whatever you have pasted, so you can compare a specific text against the baseline above. Short samples are noisy — you need several hundred letters before the pattern stabilises — but on a few pages of prose the resemblance to the table is usually striking, which is exactly why frequency analysis works.
Common questions
What is the most common letter in English?
E, by a wide margin — around 12 to 13 per cent of all letters in ordinary text. T is second at about 9 per cent, then A, O, I and N clustered between 6.5 and 8.
What is the least common letter?
Z, at roughly 0.07 per cent, with Q and J barely above it. In a 1,000-letter passage you would expect to see Z less than once.
What is the most common first letter of a word?
T, largely because of the, this, that, there and to. That differs from the overall frequency order, where E leads, because E is common inside words and rare at the start.
Why do different frequency tables disagree?
Because they were built from different corpora. A table from a dictionary counts each word once; a table from newspaper text is dominated by the most common words. Both are correct answers to different questions.
What are the most common letters in Wordle?
The same ones, adjusted for five-letter words specifically: E, A, R, O and T lead, which is why openers built from those letters perform well.
How is letter frequency used in cryptography?
It is the standard first attack on a simple substitution cipher. The most frequent symbol in the ciphertext is probably E, the most frequent three-symbol group is probably THE, and the rest follows from there.