Thai Place Names in English: What Two Sources Actually Disagree About
TL;DR: Thai place names look like chaos and mostly are not. Of the 1 364 places drawn from two sources, 721 are ones where both give the same Thai name, and of those 58% are written identically in English. A further 17.3% differ only by a class word or a space. Where the letters themselves differ, the usual cause is which words of the name a source carried over, not a rule about Thai.

What this measures
Thai is written in its own script, and everything outside Thailand has to spell it in Latin letters. Road signs, maps, passports and atlases each made that choice separately, which is why one map shows Lopburi and the next shows Lop Buri. This study asks a narrower question than whether that is confusing, and answers it across 1 364 places and 17 274 dictionary words.
Two bodies of data answer it. The first is 1 364 places from OpenStreetMap, each carrying a Thai name, an English name and a link to a Wikidata item with its own Thai and English names, which is two independent write-ups of one place. The second is a dictionary of 17 274 Thai words where every entry carries two romanisations produced under two different published systems.
Two sources agree more often than the reputation suggests
Of the 721 places where both sources have an English name and agree on the Thai, 58% are the same string. That is the finding which contradicts the folklore, and it is worth stating plainly: for most places, two people working separately reached the same Latin spelling.
58%
of Thai places carry exactly the same English name in both sources, once a disambiguating word is set aside.
The rest split into two kinds that mean very different things. A further 17.3% differ only by convention: a space one source added, a hyphen marking a syllable break, or a class word such as Airport that one source kept and the other dropped. Those are disagreements about house style, not about Thai.
| Group | Places | Share of the comparison |
|---|---|---|
| Same string | 418 | 58% |
| Convention only | 125 | 17.3% |
| Spelling differs | 178 | 24.7% |
| All compared | 721 | the whole comparison |
The last group is 24.7%, and it is the group the rest of this article is about. It is not a rounding error, and it is not the whole story either.
What the genuine differences turn out to be
Every differing pair was sorted into one category by a program, never by hand. The largest category is not a spelling rule at all: 126 of the differences are a word carried over by one source and not the other. The Thai name Nakhon Si Thammarat is written out in full by one source and run together or shortened by the other.
A tenth of the drawn places have two entirely different Thai names rather than two spellings of one name, which is 10.3% of everything that could be read. That happens because an OpenStreetMap object and a Wikidata item often describe slightly different things: the town as a place, and the municipality as a legal body containing it. 33% of the objects tagged as settlements are joined to a body rather than to the settlement.
The pairs that carry real phonetic weight, the consonants and the vowel length, are a small minority. The residual category, holding everything the instrument could not attribute, is 32 of the place differences. That is a limit on the method and not a finding about Thai, and it is reported as such.
Interpretation
The honest reading is that most of what looks like disagreement about Thai is disagreement about editing. A reader comparing two maps is far more likely to meet a dropped class word than a contested consonant.
- 126 pairs differ by a word one source carried over and the other did not. This is the largest category, and it is an editing choice.
- 10.3% of the places have two different Thai names rather than two spellings of one, usually because the two sources describe a town and its municipality.
- 178 pairs differ in the letters themselves, which is the group a reader is most likely to notice and the smallest of the three.
The two sources are less independent than they look
A study of agreement is only as good as the independence of the things being compared, and here that independence is weaker than it appears. Where a Wikidata item has an English Wikipedia article, its English label is identical to the article title 45% of the time. That is consistent with the two having a common ancestor, so part of the agreement above is one string kept in step rather than two people arriving at the same answer.
A second check is more reassuring. German labels exist for many of the same items, and a German label follows the OpenStreetMap spelling rather than the English one only 1% of the time. If OpenStreetMap were simply copying Wikipedia, the German column would be a second witness to it, and it is not.
Two published systems, one word, almost never the same answer
The second half of the question needs a different instrument, because the places give too few genuine differences to say anything reliable about mechanism. The word arm uses a Thai dictionary in which each entry carries two romanisations: one under the Paiboon system, the one most learners meet first, and one under the system the Royal Institute prescribes, which is what official signage uses. It holds 17 274 words.
Across 17 274 words the two systems agree on 1.1%. Put another way, they disagree on almost every word in the language. That is not a defect in either one. It is what happens when two systems are built for different jobs, one to be read aloud by a learner and one to be printed on a sign.
| Kind of difference | Word pairs | Share of the differences |
|---|---|---|
| How a consonant is written | 7 503 | 7 503 |
| Not one permitted alternation | 5 426 | 5 426 |
| Vowel length | 2 579 | 5 426 |
| All differences | 17 077 | every differing pair |
The largest reason is a single systematic convention. Thai separates aspirated consonants, said with a puff of air, from unaspirated ones, and English letters do not carry that distinction cleanly. The two systems made opposite choices: one writes the unaspirated series with the voiced letters, the other with the voiceless ones. Applied across a dictionary, that one decision accounts for 7 503 of the differences.
Interpretation
This is why romanisation looks chaotic and is not. The disagreement is one convention applied many times in the same direction, not many independent mistakes. A reader who knows which system a spelling comes from can predict the other spelling.

The feature that looked predictive, and was not
The study was designed to find which feature of the Thai spelling predicts where a disagreement falls. It found one, and then found the finding was an artefact. Words containing a long vowel differ between the two systems 4.7 points more often than words without one. Read alone, that looks like a mechanism.
Sorting the same words by length shows the difference shrinking as words get longer and gone among the longest. Where it goes is visible in the comparison the figure rests on rather than in the feature itself: among the shortest words the group without a long vowel differs in 80% of cases, and among the longest in 100%. A group that differs every time cannot differ more often, so at the long end there is nothing left for the feature to explain. Of the features declared before measuring, 17 274 could not be computed from published sources and only a handful varied enough to be used.
| Word length in Thai characters | Difference between the groups |
|---|---|
| Two | 20 points, interval 12.4 to 25.5 points |
| Nine or more | 0 points, interval reaching 1.5 points |
0 points
of difference remain among the longest words, and the interval around that nothing reaches 1.5 points. The feature the study set out to find was measuring how long the word is.
What our own romanisation turned out to be
Phayan shows Paiboon romanisation beside the script, and the honest finding here is not flattering. Comparing the app’s word list with the dictionary it was built from, 2 176 of 2 228 entries are identical character for character. The app is relaying a dictionary’s romanisation rather than computing its own, out of a corpus of 2 249 words.
That matters for how the app describes itself, and it is the kind of thing a product page is unlikely to volunteer. It is also why the word arm above is reported as a statement about two systems and never as agreement between two independent sources, because 2 228 of the word arm’s entries come from one dictionary.
Interpretation
A learner reading Paiboon spellings in the app is reading the same strings a paper dictionary prints, which is what those spellings are for. The limitation is that the app adds no romanisation judgment of its own.
What this does not show
The places here are ones that already carried a link between the two sources, and that link is usually added by matching the names. A place whose names disagree is therefore less likely to be in the sample at all, so 24.7% is biased downwards by an unknown amount. No correction for that is available inside this design.
Nothing here says which spelling is right. There is no authority for that and the study does not claim one: it measures where two sources part company, not which one erred. Two smaller limits are worth naming. Two of the features declared before measuring could not be computed at all. And the residual category holds 5 426 of the word differences, which is every pair the instrument could not attribute rather than a finding about Thai.
The largest single threat to the sample is not its size but its authorship. The comparison rests on 721 objects, and those came from 98 accounts, not from 721 independent decisions. The largest of them named 282 of the objects on its own.
98
accounts stand behind the 721 objects, so the effective sample is smaller than the object count suggests. The largest of them does not move the headline.
That account was checked against the rest rather than assumed harmless: its own agreement rate is 56.7%, against 58.8% for every other account together. The headline is therefore not the work of one editor. The objects were also read a second time, a day later, and all 1 364 still carried the same English name and the same Wikidata link, so the crowd did not move underneath the measurement.

How to use it
If you are reading a Thai sign and cannot match it to the map, the mismatch is most likely a class word or a space, and the table below covers the cases you will meet. Of the pairs that differ, 7 503 differ in exactly the place to look first: the unaspirated consonants.
| What you see | What it usually is |
|---|---|
| A space in one spelling and none in the other | House style. The Thai underneath is the same. |
| An extra word such as Airport or District | One source kept the class word, the other dropped it. |
| k against kh, or p against ph | The unaspirated consonants, where the systems disagree most. |
| Two entirely different names | Often two different things: a town and the municipality around it. |
Citing this study
Phayan, Thai place names in English, measured 10 Sep 2026 across 1 364 OpenStreetMap objects and 17 274 dictionary entries carrying two romanisation systems. Method, data and licence published at the dataset page.
A note on the sources
- The Royal Institute’s system states in its own documentation that it shows no tone and does not distinguish a short vowel from a long one.
- The Library of Congress table states that tonal marks are not romanized.
- The dictionary side is the Thai extract that kaikki.org publishes from English Wiktionary, where one headword can carry a Paiboon and a Royal Institute romanisation side by side. Archived in September 2026.
- OpenStreetMap data is open, and this study rests on 1364 objects that volunteers mapped and named.
What each platform says it is doing
- OpenStreetMap’s own page for the English name tag says the tag carries the name that English speakers use, and says nothing at all about how to transliterate a name written in another script. That silence is why the disagreement had to be measured rather than looked up: the tag’s rule answers a different question from the one a learner reading a sign is asking.
- Wikidata’s own help page says a label is the most common name an item would be known by, and states that it need not match the page title on a Wikimedia site. So 45% is what that documentation predicts, not a surprise in the data.
The data and the method
The measurement ran on 10 Sep 2026. Every number on this page resolves to a row in the evidence ledger, and the ledger, the data and the method are published together.
| What | Where |
|---|---|
| The dataset, one row per sample | dataset.json and dataset.csv |
| Where the rows came from, and the counts per group | metadata.json |
| The question, the counting rules and the predicted values | preregistration.md |
| The terms, which differ between the two frames | LICENCE.txt |
The place frame carries the same licence OpenStreetMap uses, and the word frame carries Wiktionary’s. No permanent identifier exists for the dataset yet; the addresses above are the stable ones.
Frequently asked questions
Is either romanisation system more correct?
Neither is wrong. Each was built for a different job: reading aloud by a learner, or printing on a sign. Phayan shows the Paiboon system beside the script and the Royal Institute form as an alternative, because a learner meeting a road sign needs the second and a learner sounding out a word needs the first. Neither is a correction of the other.
Why do maps spell the same town differently?
Usually because an editor kept a class word another dropped, or added a space between syllables. Those account for more of the differences between OpenStreetMap and Wikidata than any question about Thai sounds. Phayan teaches the class words as vocabulary rather than as noise, which is why the spelling on a sign stops looking arbitrary once the words are known.
Does this say which spelling a road sign will use?
No. The study measures where sources part company, not which spelling a given sign used. Official signage in Thailand generally follows the Royal Institute system, which Phayan offers as an alternative to Paiboon. An app that teaches one romanisation only, ThaiPod101 among them, leaves a learner unable to read the other when a sign uses it.
Can a learner use either system to read real Thai?
Both, once the mapping between them is known. The consonant differences are consistent rather than arbitrary, so a learner who has met one set can read the other after a short table. Phayan shows both beside the same word for that reason, and an app that shows only one, Ling or Drops among them, makes the other look like a different language.
Can I read the data and the method?
Yes. The question and the counting rules were fixed before the first measurement, and the dataset, the calibration fixtures, the replication report and the licence are published beside the article. Phayan publishes the record rather than a summary of it, and every number on this page resolves to a row in the evidence ledger.