
Which Thai Consonants Actually Appear, and How Often
TL;DR: Thai consonants are wildly uneven. Measured across 2,153 words weighted by how often they occur in the Thai National Corpus, the top fifteen account for 83 percent of every consonant you will meet, and two of the forty-four do not occur at all. But knowing 83 percent of the letters does not mean reading 83 percent of the words, because one unknown letter stops a whole word. On that measure fifteen gets you 72 percent, and the curve does not turn useful until about twenty.

What was counted, and in what
Phayan ships a word list of 2,249 Thai words, and 2,153 of them carry a frequency count from the Thai National Corpus by way of pythainlp. That count is the number of times the word appears in a large body of real written Thai, and it is the thing that makes this measurable rather than impressionistic. Without it you can only count how many words in a dictionary contain a letter, which treats ฒ in ผู้เฒ่า as equal to น in ใน.
So every consonant in every word was counted once per occurrence of that word. That gives 35,006,713 weighted consonant occurrences, which is a proxy for the consonants a reader actually passes their eyes over. The letter inventory is Phayan's own: forty-four symbols, of which ฃ kho khuat and ฅ kho khon are obsolete, leaving 42 in everyday use. That framing is covered properly in the beginner's overview of the script, and this post is only about how often each one turns up.
Two honest limits before the numbers. The word list is a curated learner corpus rather than a slice of the newspaper, so it leans toward the common and the concrete. And frequency is counted over words, not over the running text of any particular thing you might read; a Thai legal notice and a Thai menu have different distributions from each other and from this.
The first fifteen letters and what they buy
Here is the top of the list, with each letter's share of all weighted consonant occurrences and the running total.
| # | Letter | Name | Share | Running total |
|---|---|---|---|---|
| 1 | น | no nu | 11.3% | 11.3% |
| 2 | อ | o ang | 9.1% | 20.4% |
| 3 | ก | ko kai | 8.3% | 28.7% |
| 4 | ง | ngo ngu | 7.4% | 36.1% |
| 5 | ม | mo ma | 7.3% | 43.4% |
| 6 | ร | ro ruea | 5.9% | 49.3% |
| 7 | ว | wo waen | 4.9% | 54.2% |
| 8 | ด | do dek | 4.1% | 58.2% |
| 9 | ย | yo yak | 4.0% | 62.3% |
| 10 | ท | tho thahan | 4.0% | 66.2% |
| 11 | ล | lo ling | 3.9% | 70.1% |
| 12 | ห | ho hip | 3.5% | 73.6% |
| 13 | ป | po pla | 3.3% | 76.9% |
| 14 | ค | kho khwai | 3.2% | 80.1% |
| 15 | ข | kho khai | 3.1% | 83.2% |
Five letters get you to 43 percent. Ten get you to 66. The next five are worth roughly three points each and take you to 83, and then the returns fall off a cliff: จ at number sixteen is still worth 3.1 percent, but ส at nineteen is worth 1.5, ช at twenty-one is worth 1.0, and everything from number twenty-five down is worth less than three tenths of a percent each.

Covering the letters is not covering the words
This is the part that surprised the person who ran the numbers, and it is the reason this post is not just a table. A percentage of letters is not a percentage of reading, because a word is only readable when you know every consonant in it. One unfamiliar letter in the middle of a five-letter word costs you the whole word.
So the same corpus was counted again, this time asking a different question: with only the top N consonants, how much running text consists of words you can decode outright?
| Letters known | Share of consonants seen | Share of running words fully readable |
|---|---|---|
| 5 | 43.4% | 20.5% |
| 10 | 66.2% | 45.3% |
| 15 | 83.2% | 72.4% |
| 18 | 91.7% | 86.1% |
| 20 | 94.7% | 91.0% |
| 25 | 98.4% | 97.4% |
| 30 | 99.4% | 99.0% |
The gap between the two columns is the cost of the word rule, and it is widest exactly where a beginner is. At five letters you have seen nearly half the consonants in front of you and can read a fifth of the words. At ten you have seen two thirds and can read under half. The two columns only converge around twenty, which is roughly where reading stops feeling like decoding and starts feeling like reading.
That is the honest answer to how long the alphabet takes. It is not fifteen letters and you are away. It is about twenty before the letters stop being the bottleneck, and everything after that is vowels, tone rules and vocabulary. The shape of that longer curve is in how long it takes to learn to read Thai.
The nine letters you will almost never have to decode
At the other end of the list the numbers get very small indeed. These nine letters, taken together, account for about six hundredths of one percent of the consonants in the corpus.
- ฮ ho nok-huk and ฆ kho ra-khang are each in fewer than twenty of the 2,153 words.
- ฎ do cha-da and ฏ to pa-tak appear in three words and six words respectively, mostly Sanskrit and Pali borrowings and royal or formal vocabulary.
- ฬ lo chu-la, ฌ cho choe, ฑ tho montho and ฒ tho phu-thao appear in five, two, two and two words each. You will meet ฬ in จุฬา and ฒ in ผู้เฒ่า and that may be most of your encounters for a year.
- ฃ kho khuat and ฅ kho khon appear in zero words. They are the obsolete pair, retired from Thai typewriters in the 1960s, and the corpus confirms it: not one occurrence.
None of that means skip them. It means do not spend week one on them, and do not let a chart that presents forty-four equal cells convince you that they are equally worth your Tuesday evening.

Frequency is not the same as a good order to learn in
It would be neat if the answer were to learn the letters in the order of the first table, and it is not. Two things get in the way.
The first is that Thai consonants carry a class, and the class is half of how the tone of a syllable is decided. Learning น ง ม ร ว ย ล in a block teaches you seven letters that are all low class, which means you can read them and still have no way to work out what tone the syllable takes. A teaching order has to hand you enough of a class to make the tone rule usable, which is why the class system shapes the sequence more than frequency does.
The second is that letters are not learned in isolation; they are learned inside words, and a word needs a vowel. The top five consonants build almost nothing on their own. Phayan sequences its levels around what makes a readable word at each step rather than around raw counts, and audio on every word is what stops the letters from becoming shapes with no sound attached.
The honest use of this table is not as a syllabus. It is as a permission slip: when you are on letter nineteen and the chart still shows twenty-five to go, the remaining twenty-five are worth about 5 percent of what you will read, and the hard part is already behind you. The comparison of what each learning method skips makes the same argument about alphabet order from the other direction.
What to do with this on a Tuesday
Three things follow from the numbers and nothing else does.
- If you are picking flashcards by hand, weight the first twenty. Beyond that you are optimising the last 5 percent of your reading.
- If a word on a sign has one letter you do not know, that is normal and it is the word rule doing its thing rather than a gap in your study. Look it up, move on.
- Do not measure your progress against forty-four. Measure it against twenty, then against the vowels, which are a bigger job than the consonants ever were.
The pebbles in the photograph at the top are laid out in the shape of that first table. That is what the Thai alphabet looks like when you count it instead of listing it, and it is a much friendlier shape than a grid of equal squares.
Frequently asked questions
What is the best order to learn the Thai consonants in?
An order that groups them by consonant class while still front-loading the common ones, rather than either pure frequency or the traditional chart order. Phayan sequences its levels that way, introducing enough of one class at a time that the tone rule becomes usable instead of theoretical. Anki with a frequency-ordered deck is the closest do-it-yourself equivalent, and it will teach the letters faster while leaving the tone rules entirely to you.
How many Thai consonants do I need before I can read signs?
About twenty, on this measurement: twenty consonants make 91 percent of running words fully decodable, where fifteen only manages 72 percent. Phayan reaches that point across its first few levels, and the free levels cover the five tones, first words, more vowels and the finals. A sign is easier than the average because place names and shop words repeat, so real-world reading tends to arrive a little before the number suggests.
Is it safe to skip the rare Thai letters at the start?
Yes, for a while, as long as skip means postpone. The nine rarest consonants together account for roughly six hundredths of one percent of the consonants in this corpus, and two of them do not appear at all. Phayan still teaches every letter in use, because the rare ones cluster in formal and royal vocabulary you will eventually meet on a document or a temple sign, and a letter you have never seen stops a word dead.
What is the difference between letter frequency and word coverage?
Letter frequency counts how often a symbol appears; word coverage counts how much text you can actually read, and a word only counts when every consonant in it is known. That is why fifteen letters cover 83 percent of Thai consonants but only 72 percent of running words. Phayan tracks the second one, because it is the number a learner feels. ThaiPod101 and other audio-first courses sidestep the question by not asking you to decode at all.
Why does every Thai alphabet chart show forty-four if two are dead?
Because the chart is an inventory of the writing system rather than a study plan, and ฃ and ฅ are still part of the system even though they were dropped from typewriters in the 1960s and appear in none of the 2,153 words counted here. Phayan lists them for completeness and does not drill them. If a paper chart is what you like, print one and cross those two out.