Written when we measure something worth writing down. Every figure here comes from the same engine the tools run on.
o200k_base was supposed to close the gap for non-English text. Running both tokenizers over a parallel corpus in 25 languages shows it closed most of it for some languages and almost none for others — and reshuffled which languages are now the most expensive.