The price of a language
The same sentence can produce up to 15 times more tokens depending on the language it is written in. This is no matter of style: it is an engineering choice, with measurable consequences.
Putting the same question to a language model, in two different languages, does not cost the same to compute. It is not a matter of the length of words or of politeness: it is a measured, documented gap, and it comes from an engineering choice made long before anyone asks a question.
What a text becomes before it enters the model
A language model does not read words. It reads tokens: fragments of text, shorter than a word or sometimes longer, cut up by a program called a tokeniser before the text enters the calculation. The detail of that cutting comes down to a fixed vocabulary, learnt once and for all on a very large corpus of texts. Every word has a number of tokens of its own, and that number is what the model charges for, in computing time as much as in room in its working memory.
The same content, very different totals
A 2023 study measured this gap on a set of tokenisers genuinely used by language models. Its finding: the same content, translated into different languages, can produce up to 15 times more tokens in the least favoured language than in the most favoured one. Even the tokenisers that cut at the level of the character or the byte, supposed to avoid this problem, keep a gap measured at more than 4 times over some pairs of languages.
15×
the token gap measured for the same content, between the most favoured and the least favoured language, in the worst cases (Petrov et al., 2023).
More tokens to say the same thing is no aesthetic detail. It is more computing at every answer, so slower to generate and more expensive to run for a service billed by the token. It is also a context window, the limited working memory of a model, that fills up faster: for the same number of tokens allowed, a text in a disfavoured language holds less genuinely useful content.
Why this choice is not neutral
The most widespread cutting method is called BPE, byte pair encoding. It was not invented for language. Philip Gage described it in 1994 as a general method for compressing computer files: it spots the most frequent pair of bytes in some data, replaces it with an unused byte, and repeats the operation until there is nothing left to gain. Nothing to do, originally, with a human language.
In 2016, Rico Sennrich, Barry Haddow and Alexandra Birch took the idea up for a precise problem in machine translation: what to do with a rare word, absent from the vocabulary a model has learnt? Their answer cuts the word into more common fragments, already known to the model. The method works well, and it has since been taken up by very nearly every large language model.
A sub-word vocabulary is learnt on a corpus, once, before the training of the model itself. A mostly English-language corpus learns a vocabulary optimised for English: the most common fragments in that language become single, short, few tokens. A word from a language less represented in that corpus ends up cut into more fragments, simply because its own frequent repetitions did not have the same chance of being seen while the vocabulary was being learnt. This is no property of language. It is the imprint of a corpus.
What this changes in practice
A service billed by the token costs more, for equivalent content, to whoever writes in a language disfavoured by the tokeniser. The answer also takes longer to appear, since a model produces its tokens one after another, at a roughly constant speed: more tokens to produce, more waiting. And the context window, that fixed number of tokens a model can read or write at once, holds less useful text: a document, a conversation history, a long instruction all fit into it less well.
None of these effects comes from one language being more complex than another. They come from a vocabulary learnt once, on a given corpus, never recalculated according to who uses it afterwards. The authors of the 2023 study put forward one lead: designing, from the outset, tokenisers that do not structurally favour one language over another.
Sources
- Language Model Tokenizers Introduce Unfairness Between Languages Aleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel Bibi, 2023.
- Neural Machine Translation of Rare Words with Subword Units Rico Sennrich, Barry Haddow, Alexandra Birch, 2016.
- A New Algorithm for Data Compression Philip Gage, 1994.
This article is published under the CC BY 4.0 licence: copy it, translate it, republish it, crediting ODERSA.