This is an official ODERSA website. Here’s how you know

The official domain

The address of this site ends in odersa.org. Every service the association runs sits on a subdomain of odersa.org and nowhere else. If the address in your browser’s address bar does not end in odersa.org, this site is not ours.

Free, and no account

Everything is open straight away. No sign-up, no account, no password, no subscription, no advertising. Nothing is held back for those who pay, because there is nothing to pay for.

No data collected

This site does not follow you: no tracker, no tracking cookie, no measurement tool built into these pages, and nothing measured on your device. Our host counts requests in aggregate, as any server that answers does: a total, never a profile. You do not have to take our word for it: open your browser’s developer tools, go to the Network tab, and reload the page. You will see the full list of what the site asks for. Everything comes from odersa.org, nothing goes anywhere else.

Free to reuse

The content is published under the CC BY 4.0 licence. You may copy it, translate it, print it and pass it on, for your classes as much as for the people around you, on one condition only: credit ODERSA.

The price of a language

The same sentence can produce up to 15 times more tokens depending on the language it is written in. This is no matter of style: it is an engineering choice, with measurable consequences.

Putting the same question to a language model, in two different languages, does not cost the same to compute. It is not a matter of the length of words or of politeness: it is a measured, documented gap, and it comes from an engineering choice made long before anyone asks a question.

What a text becomes before it enters the model

A language model does not read words. It reads tokens: fragments of text, shorter than a word or sometimes longer, cut up by a program called a tokeniser before the text enters the calculation. The detail of that cutting comes down to a fixed vocabulary, learnt once and for all on a very large corpus of texts. Every word has a number of tokens of its own, and that number is what the model charges for, in computing time as much as in room in its working memory.

The same content, very different totals

A 2023 study measured this gap on a set of tokenisers genuinely used by language models. Its finding: the same content, translated into different languages, can produce up to 15 times more tokens in the least favoured language than in the most favoured one. Even the tokenisers that cut at the level of the character or the byte, supposed to avoid this problem, keep a gap measured at more than 4 times over some pairs of languages.

15×

the token gap measured for the same content, between the most favoured and the least favoured language, in the worst cases (Petrov et al., 2023).

More tokens to say the same thing is no aesthetic detail. It is more computing at every answer, so slower to generate and more expensive to run for a service billed by the token. It is also a context window, the limited working memory of a model, that fills up faster: for the same number of tokens allowed, a text in a disfavoured language holds less genuinely useful content.

Why this choice is not neutral

The most widespread cutting method is called BPE, byte pair encoding. It was not invented for language. Philip Gage described it in 1994 as a general method for compressing computer files: it spots the most frequent pair of bytes in some data, replaces it with an unused byte, and repeats the operation until there is nothing left to gain. Nothing to do, originally, with a human language.

In 2016, Rico Sennrich, Barry Haddow and Alexandra Birch took the idea up for a precise problem in machine translation: what to do with a rare word, absent from the vocabulary a model has learnt? Their answer cuts the word into more common fragments, already known to the model. The method works well, and it has since been taken up by very nearly every large language model.

A sub-word vocabulary is learnt on a corpus, once, before the training of the model itself. A mostly English-language corpus learns a vocabulary optimised for English: the most common fragments in that language become single, short, few tokens. A word from a language less represented in that corpus ends up cut into more fragments, simply because its own frequent repetitions did not have the same chance of being seen while the vocabulary was being learnt. This is no property of language. It is the imprint of a corpus.

What this changes in practice

A service billed by the token costs more, for equivalent content, to whoever writes in a language disfavoured by the tokeniser. The answer also takes longer to appear, since a model produces its tokens one after another, at a roughly constant speed: more tokens to produce, more waiting. And the context window, that fixed number of tokens a model can read or write at once, holds less useful text: a document, a conversation history, a long instruction all fit into it less well.

None of these effects comes from one language being more complex than another. They come from a vocabulary learnt once, on a given corpus, never recalculated according to who uses it afterwards. The authors of the 2023 study put forward one lead: designing, from the outset, tokenisers that do not structurally favour one language over another.

Sources

This article is published under the CC BY 4.0 licence: copy it, translate it, republish it, crediting ODERSA.

Back to the blog