Letter frequency in English
English does not use its letters evenly, and it is not close. E turns up roughly a hundred and eighty times as often as Z. That imbalance is the reason a substitution cipher can be solved at all, it is why every word game you have ever played is quietly built around the same nine letters, and it is one of the oldest measured facts about the language — it was written down in the ninth century, long before anyone had a word for statistics.
Here are the numbers, and then the part most pages leave out: when they stop being true.
The table
Percentages are shares of all letters in ordinary English prose, from the standard reference count used throughout cryptography teaching. They will not match any single sentence and are not supposed to.
| Letter | % |
|---|---|
| E | 12.7% |
| T | 9.06% |
| A | 8.17% |
| O | 7.51% |
| I | 6.97% |
| N | 6.75% |
| S | 6.33% |
| H | 6.09% |
| R | 5.99% |
| D | 4.25% |
| L | 4.03% |
| C | 2.78% |
| U | 2.76% |
| M | 2.41% |
| W | 2.36% |
| F | 2.23% |
| G | 2.02% |
| Y | 1.97% |
| P | 1.93% |
| B | 1.29% |
| V | 0.98% |
| K | 0.77% |
| J | 0.15% |
| X | 0.15% |
| Q | 0.1% |
| Z | 0.07% |
Three numbers are worth carrying in your head. The top nine letters — E T A O I N S H R — are about seventy per cent of all English text. The top twelve are about eighty. And E alone is one letter in eight, which is why it is the first guess on any board and why we never hand it to you free.
Different corpora move the tail around. Count a news archive, a novel and a pile of scientific papers and you will get slightly different figures — H and R trade places, C creeps above U, and any text about zebras or quartz will embarrass the bottom row. The first six are stable across all of them. Trust the head of this table; hold the tail loosely.
ETAOIN SHRDLU, and what the mnemonic gets wrong
ETAOIN SHRDLU is how most people remember the order, and it is a printers' artefact rather than a linguistic finding: it is the first two columns of a Linotype machine's keyboard, which the nineteenth-century typesetters arranged by how often they reached for each letter. It is a remarkably good approximation for something derived from muscle memory.
It is not exact. Modern counts put H above R, and put C above U — so the true order runs closer to E T A O I N S H R D L C U. The mnemonic is worth keeping for the first six letters and worth distrusting after that, which is more or less the theme of this entire page.
Why the table breaks on short text
Here is the thing that gets skipped, and it is the difference between using frequency well and losing puzzles to it.
The table describes millions of letters. A cryptogram is thirty to a hundred. At that size you are not measuring English, you are measuring one sentence — and one sentence has opinions. A single uncommon word distorts the whole count. A quotation that happens to repeat a word twice doubles four or five letters at once. E's expected share is 12.7%, but on a forty-letter sample the ordinary random spread around that is wide enough to put E anywhere from first to fourth without anything unusual having happened at all.
Rough working numbers, and they are the honest answer to "when does counting start working?":
- Under about 40 letters (roughly 8–10 words). The count is close to noise. The top symbol is as likely to be T, A or O as E. Solve on word shape and treat the count as a tiebreak only.
- Around 70–100 letters (15–20 words). E and T usually settle into the top two. The tier below — A, O, I, N — is still shuffled and should not be committed to on count alone.
- Above roughly 200 letters. The top nine start resembling the table in something like the right order. Almost no cryptogram is this long.
Which means: on the puzzle in front of you, frequency is a prior, not a proof. Everything on this page is a ranked hunch that structure is allowed to overrule.
What is stable at puzzle length, and what isn't
Not all of the table degrades at the same rate. At forty to eighty letters:
- Reliable.Something in the top group is frequent — you will not find a board where E, T, A, O, I, N, S, H and R are all rare, because they are seventy per cent of the language. Rare symbols are genuinely rare letters: a symbol appearing once or twice on a long board is J, K, Q, V, X or Z far more often than chance suggests.
- Unreliable. The exact ranking inside the top group, and any comparison between two symbols whose counts differ by one or two. A symbol on six occurrences and a symbol on five tell you nothing about each other.
The practical rule that falls out: use the count to pick a set of candidates, never a single letter, and let word shape choose from the set.
Position beats frequency more often than you would think
Where a letter sits is often better evidence than how often it appears, and it costs nothing extra to look.
- Word-final. E ends more English words than any other letter, by a distance; then S, D, T and N. A symbol that is frequent and keeps landing last is E with high confidence — much higher than the raw count alone would justify.
- Word-initial. T, A, O, S and W lead. E is conspicuously scarce here relative to how common it is overall, which is the single most useful asymmetry in the language: a frequent symbol that never opens a word is a strong E, and a frequent symbol that opens three words probably is not.
- One-letter words. A or I, with no third option.
Pairs, doubles and the counts that actually pay
Letters do not arrive independently, and the pair statistics survive a substitution cipher exactly as the single-letter ones do.
- The most common pairs are TH, HE, IN, ER, AN, RE, ND. TH and HE together are why a three-symbol word is so often THE and why cracking it is worth three letters rather than one.
- The most common doubles are LL, EE, SS, OO and TT, then FF, MM, NN, RR and PP. A doubled symbol is therefore a five-way guess before any other evidence, and usually a two-way one after.
- Q is followed by U, without exception in ordinary English.
Those are the counts with the best yield per second spent, and none of them require you to tally anything — see the full pattern reference for how they combine with word shape.
A real board where the top letter is not E
Why this board. Thirty-four letters, and the most frequent one is not E. O appears eight times — 23.5% of the board, three times its expected share — while E appears three times and comes joint third. A solver who applies "most frequent equals E" mechanically loses the board on move one. This is the page's whole argument, on a real puzzle, with real counts.
| Symbol | 1 | 10 | 4 | 13 | 6 | 8 | 15 | 17 | 19 | 21 | 2 | 23 | 25 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Count | 8 | 5 | 3 | 3 | 2 | 2 | 2 | 2 | 2 | 2 | 1 | 1 | 1 |
| Share | 23.5% | 14.7% | 8.8% | 8.8% | 5.9% | 5.9% | 5.9% | 5.9% | 5.9% | 5.9% | 2.9% | 2.9% | 2.9% |
- Count first — that part of the advice is right. Thirty-four letters, thirteen distinct symbols. Symbol 1 appears eight times, symbol 10 five times, and everything else three or fewer.
- Now check the count against the table before you spend it. Symbol 1 is at 23.5%. E's share in English is 12.70%, and no letter in ordinary English text exceeds it. A symbol sitting at nearly double the language's maximum is not emphatically E — it is a signal that thirty-four letters is far too small a sample for the table to mean much. That thought is worth five seconds and it saves the board.
- Position settles it. Symbol 1 opens a five-symbol word (
1 17 13 6 10), closes a three-symbol word (25 4 1), and appears doubled inside two different words (10 8 4 1 1 15and2 1 1 19). E in English is heavily word-final and conspicuously rare word-initial. A letter that is common in first position, last position and as a double is far better explained by O. Commit 1 = O — on positional evidence, against the raw count. - Spend the given letters. A = 21 and N = 6, everywhere they occur.
- The three-letter word ending in O.
25 4 1is? ? O. The candidates are few — WHO, TWO, AGO, TOO. And symbol 4 is also the first letter of the sentence's opening two-letter word4 13. Only WHO gives a live two-letter word: 25 = W, 4 = H, so4 13isH ?→ HE, giving 13 = E. - Look at what just happened. E is symbol 13, on three occurrences — joint third on this board, behind a symbol on eight and one on five. The most common letter in the English language came third in a nine-word sentence, and nothing unusual caused it. That is the sample size, not the quote.
- The five-letter word.
1 17 13 6 10isO ? E N ?→ OPENS: 17 = P, 10 = S. Note symbol 10 was the second most frequent at five, and it is S — which the table ranks seventh. The count was roughly right about that one. It is right often enough to be worth doing and wrong often enough to need checking. - The first doubled O.
10 8 4 1 1 15isS ? H O O ?→ SCHOOL: 8 = C, 15 = L. - The second.
2 1 1 19is? O O ?and follows "school": DOOR, giving 2 = D, 19 = R. - Forced.
8 15 1 10 13 10is alreadyC L O S E S. No guess required. - Last word.
17 19 23 10 1 6isP R ? S O N→ PRISON: 23 = I.
Read it:"He who opens a school door, closes a prison."
The count was consulted twice and overruled once. That is the correct ratio, and it is why every other page on this site calls frequency a hint rather than a rule.
Or browse more Victor Hugo cryptograms.
The contrast, on a longer real board."Whenever you find yourself on the side of the majority, it is time to pause and reflect." — Mark Twain, this seventy-letter board, tier 4, Hard. Seventy letters, and the counts come out E 11, T 7, O 6, I 6 — 15.7%, 10.0%, 8.6%, 8.6% against the table's 12.70, 9.06, 7.51, 6.97. The top two land exactly where the table says they will. The tier below still does not: A, which the table ranks third at 8.17%, appears three times on that board — 4.3%, roughly half its expected share. Thirty-four letters was too few for the first place. Seventy is enough for the first two places and no more. That is the degradation curve, measured on our own puzzles rather than asserted.
Go and count something
The fastest way to feel how noisy a short count is: try it on medium cryptograms, then on hard cryptograms, which are longer and where the numbers start behaving. The method these numbers slot into is on how to solve cryptograms.
Common questions
What are the most common letters in English?
In order: E, T, A, O, I, N, S, H, R — E at about 12.7% of all letters, T at 9.1%, A at 8.2%, down to R at 6.0%. Those nine make up roughly seventy per cent of ordinary English text, and the top twelve about eighty per cent. The exact percentages shift a little depending on what is being counted — fiction, news and technical writing all differ — but the first six letters are stable across every corpus anyone has measured.
What does ETAOIN SHRDLU mean?
It is a mnemonic for the twelve most common letters of English, in descending order of frequency, and it comes from Linotype typesetting machines — the first two columns of the keyboard were arranged by how often printers reached for each letter, so running a finger down them produced "etaoin shrdlu". It is a good approximation rather than an exact ranking: modern counts put H above R, and C above U. Treat the first six as reliable and the rest as a rough guide.
What is the least common letter in English?
Z, at about 0.07% of letters — roughly one in fourteen hundred. Q, J and X are the other three below a fifth of a percent. This is a count of letters as they appear in running text, not of dictionary entries; by number of words the order differs, since a rare letter can appear in many rarely-used words.
Does letter frequency work on short texts?
Only loosely, and this is the caveat most guides skip. The percentages describe millions of letters; a cryptogram is thirty to a hundred. Under about forty letters — eight to ten words — the count is close to noise and the most frequent symbol is roughly as likely to be T, A or O as E. Around seventy to a hundred letters, E and T usually settle into the top two, but the tier below them is still shuffled. On a real puzzle, use the count to narrow candidates and let word shape make the final choice.
Which letters are most common at the start and end of words?
Word-initial position is dominated by T, A, O, S and W. Word-final position by E, S, D, T and N. The useful asymmetry is E: it is the most common letter overall and the most common final letter, but it is conspicuously scarce as a first letter. So a frequent symbol that keeps landing at the end of words is almost certainly E, while a frequent symbol that opens several words probably is not — positional evidence that is often stronger than the raw count.
How do you use letter frequency to solve a cryptogram?
Count the symbols, take the most frequent one and try it as E, then work through T, A, O, I, N, S, R and H for the next tier — but check each candidate everywhere it appears before committing, and never let the count override a word shape. In practice frequency is the third move, not the first: on a short puzzle the one- and two-letter words and the repeated words are more reliable evidence. The full order of operations is on how to solve cryptograms.