Did the new tokenizer fix multilingual AI costs? We measured it.
o200k_base was supposed to close the gap for non-English text. Running both tokenizers over a parallel corpus in 25 languages shows it closed most of it for some languages and almost none for others — and reshuffled which languages are now the most expensive.
When OpenAI shipped o200k_base with GPT-4o, one of the stated improvements was better handling of non-English text. That mattered commercially, not just technically: providers bill per token, so a tokenizer that splits your language into more pieces charges you more for saying the same thing.
The claim was easy to make and rarely checked. So we checked it. Five texts in five registers, translated into 25 languages, run through both cl100k_base and o200k_base. Every number below is counted, not estimated.
The short answer: yes, but unevenly
The new tokenizer did close the gap — dramatically for some languages and barely at all for others. Hindi went from 5.35× English to 1.67×, erasing 85% of its penalty. Hungarian moved from 2.40× to 2.02×, erasing 27%.
The uneven improvement matters more than the average one, because it changed the ranking. Hungarian was the eighth most expensive language under cl100k. Under o200k it is second.
| Language | GPT-4 | GPT-4o | Gap closed |
|---|---|---|---|
| Hindi | 5.35× | 1.67× | 85% |
| Hebrew | 3.89× | 1.58× | 80% |
| Arabic | 3.18× | 1.45× | 79% |
| Chinese | 1.64× | 1.16× | 75% |
| Greek | 5.42× | 2.32× | 70% |
| Dutch | 1.63× | 1.19× | 70% |
| Russian | 2.55× | 1.49× | 68% |
| Vietnamese | 2.62× | 1.54× | 67% |
| Korean | 2.43× | 1.50× | 65% |
| Ukrainian | 3.16× | 1.91× | 58% |
This table is generated from live measurements, not copied into the text.
Why some languages were left behind
The languages that improved most were those paying a byte-level penalty: Hindi, Hebrew and Arabic use scripts that cost multiple UTF-8 bytes per character, and a larger vocabulary gives those scripts far more merged pieces to work with. That is a problem you can fix by spending vocabulary on it.
The languages that improved least were those paying a morphological penalty. Hungarian, Finnish and Turkish are agglutinative: they build long words by stacking suffixes, producing word forms so numerous that no fixed vocabulary can cover them. A bigger vocabulary helps at the margin, but it cannot enumerate a combinatorial space.
Byte problems shrink when you buy more vocabulary. Morphology problems do not.
What this costs in practice
For a product handling 500,000 calls a month at 500 English tokens each, priced at $2.50 per million input tokens, the language penalty alone runs from roughly $1,200 a year for Chinese to nearly $10,000 for Greek. None of that buys a better answer.
There is also a version-lag problem that is easy to miss. If your product still runs on a GPT-4 generation model, you are billed at the old tokenizer's rate regardless of what newer models can do. For Turkish that is 2.17× rather than 1.64×; for Greek, 5.42× rather than 2.32×. Migrating models is a pricing decision, not only a capability one.
The full ranking
| Language | GPT-4 | GPT-4o |
|---|---|---|
| Greek | 5.42× | 2.32× |
| Hungarian | 2.40× | 2.02× |
| Ukrainian | 3.16× | 1.91× |
| Czech | 2.42× | 1.86× |
| Polish | 2.12× | 1.78× |
| Romanian | 1.97× | 1.68× |
| Hindi | 5.35× | 1.67× |
| Turkish | 2.17× | 1.64× |
| Japanese | 2.28× | 1.63× |
| Finnish | 2.07× | 1.62× |
| Hebrew | 3.89× | 1.58× |
| Vietnamese | 2.62× | 1.54× |
| Korean | 2.43× | 1.50× |
| Italian | 1.70× | 1.49× |
| Russian | 2.55× | 1.49× |
| Arabic | 3.18× | 1.45× |
| German | 1.70× | 1.38× |
| Indonesian | 1.66× | 1.36× |
| Swedish | 1.66× | 1.35× |
| Portuguese | 1.55× | 1.26× |
| Spanish | 1.45× | 1.25× |
| French | 1.51× | 1.24× |
| Dutch | 1.63× | 1.19× |
| Chinese | 1.64× | 1.16× |
This table is generated from live measurements, not copied into the text.
Method and its limits
Both tokenizers were run over a parallel corpus of five texts per language, covering business email, support reply, product description, instructions and casual chat. Token counts are exact. The corpus is published in full in the repository.
The honest caveat is translation. These are our translations, and a wordier rendering costs more tokens, so the corpus is auditable by design. We also publish the spread across the five registers rather than a single average — Turkish, for instance, ranges from 1.18× on casual chat to 2.09× on formal writing, and quoting only the flattering end of that would be a choice, not a measurement.