Tokens

The same sentence costs more in Arabic

Every model is priced per token, and every price list is easy to compare. What no price list tells you is how many tokens your text actually becomes — and that depends on the script you write in. The same meaning can cost 429 tokens in Arabic on one model and 1483 on another, for identical work at an identical price per token. These are measured, not estimated.

Arabic · Model1.19×ALLaM 7B
Arabic · Model4.22×SmolLM2 1.7B
What it meansA 3.6× difference nobody quotes

Your bill is the price per token multiplied by the number of tokens. The first half is published everywhere; the second half is not published at all. On an Arabic workload the second half varies by 3.6× across the models measured here, which can outweigh any price difference between them.

See it happen

One sentence, eighteen ways

The same words through every tokenizer, so the only thing changing between these numbers is which vocabulary is reading them. English lands on about twenty tokens almost everywhere. The Arabic runs from 13 tokens to 52 — 4.0× apart, for one sentence nobody edited in between. A model that learned the script keeps whole words; one that didn't spends several tokens per letter.

Illustration only — excluded from the benchmark

In the UAE, everyone is Emirati through their love for this land and their contributions to it.

في الإمارات الكل إماراتي، بحبه لهذه الأرض وعطائه لها.

HH Sheikh Mohamed bin Zayed Al Nahyan, President of the UAEMar 2026Quoted from

ALLaM 7B

SDAIALLMVocabulary 64,000
English21 tokens
IntheUAE,everyoneisEmiratithroughtheirloveforthislandandtheircontributionstoit.
العربية13 tokens
فيالإماراتالكلإماراتي،بحبهلهذهالأرضوعطائهلها.

A box marked · is a token that isn't a character at all — a fragment of a UTF-8 sequence that means nothing on its own. They are shown rather than hidden, because a row of them is the finding.

Note that this Arabic is shorter than its English on the better tokenizers. Single sentences vary, and this one is unusually compact in Arabic — the figures measured across the whole corpus, where Arabic costs more on every tokenizer without exception, are in the table below.

Large or small

Does a small model cost less?

A small model is the usual advice for cutting cost: it runs on your own hardware, it is cheaper per token, and for many tasks it is good enough. On an Arabic workload that advice is incomplete, because a small model is often given a small vocabulary — and a small vocabulary is paid for by whichever script it was not built around.

Large models14
1.19× to 3.83× on Arabic
Small models5
1.32× to 4.22× on Arabic

But small does not mean bad, and large does not mean safe. Across the models measured here the two groups cover almost exactly the same ground: large models run from 1.19× to 3.83×, small ones from 1.32× to 4.22×. The best small model beats all but one of the large ones. The worst large model is worse than most of the small ones.

The clearest proof is a pair from one family. These score identically, to the second decimal, because they share a tokenizer — so on Arabic the smaller one carries no penalty at all for being smaller:

Small modelsGemma 2 2B256,0001.67×identicalLarge modelsGemma 2 9B256,000
Small modelsQwen2.5 0.5B151,6431.81×identicalLarge modelsQwen3 8B151,643

Ask what vocabulary a model was given, not how many parameters it has. The first tells you what your Arabic will cost; the second tells you nothing about it.

Every tokenizer measured

Arabic and Chinese, against English

Tokens needed for the same 6 articles, relative to English. Below 1.00 means the language is cheaper than English on that model. Sorted by the Arabic figure, best first.

Measured corpus — 6 UN-hosted articles

ModelVocabularyEnglishArabicChinese
ALLaM 7BSDAIALLM64,0003621.19×3.87×
Falcon H1 34BTIILLM261,1203501.31×0.87×
BLOOMZ 560MBigScienceSLM250,6803561.32×0.83×
Aya 101CohereLLM250,1004631.35×0.80×
Jais family 13BInceptionLLM84,9923561.37×2.11×
GPT-4o / GPT-5 familyOpenAILLM—3521.62×1.16×
gpt-oss 20BOpenAILLM199,9983521.62×1.16×
Fanar 1 9BQCRILLM128,2563571.63×1.77×
Gemma 2 9BGoogleLLM256,0003501.67×1.01×
Gemma 2 2BGoogleSLM256,0003501.67×1.01×
DeepSeek V3DeepSeekLLM128,0003511.74×0.90×
Qwen3 8BAlibabaLLM151,6433541.81×0.92×
Qwen2.5 0.5BAlibabaSLM151,6433541.81×0.92×
Llama 3.xMetaLLM128,0003541.82×1.22×
AceGPT v2 8BFreedomIntelligenceLLM128,0003541.82×1.22×
GPT-4 / GPT-3.5OpenAILLM—3543.25×1.76×
Phi-3.5 miniMicrosoftSLM32,0003993.60×1.88×
Mistral 7BMistral AILLM32,7683693.83×1.54×
SmolLM2 1.7BHugging FaceSLM49,1523514.22×2.52×
Why

It's the vocabulary, not the size

A tokenizer has a fixed vocabulary, and everything outside it gets broken into pieces. Models with a large vocabulary have room for Arabic and Chinese words; models with a small one spend that room on English and pay for every other script by the letter. This tracks vocabulary size almost exactly — and not model size. BLOOMZ 560M has 250,680 tokens in its vocabulary and handles Arabic better than models many times its size, while Mistral 7B does not. Two models in the same family, one sixteen times the other, score identically because they share a tokenizer.

Vocabulary sizeArabic tokens per English token
Falcon H1 34B1.31×
Gemma 2 9B1.67×
Gemma 2 2B1.67×
BLOOMZ 560M1.32×
Aya 1011.35×
gpt-oss 20B1.62×
Qwen3 8B1.81×
Qwen2.5 0.5B1.81×
Fanar 1 9B1.63×
DeepSeek V31.74×
Llama 3.x1.82×
AceGPT v2 8B1.82×
Jais family 13B1.37×
ALLaM 7B1.19×
SmolLM2 1.7B4.22×
Mistral 7B3.83×
Phi-3.5 mini3.60×
What it costs

Monthly and yearly, on your own numbers

Set your workload and the price you actually pay. The token counts are measured; the price is yours, because the models with published prices and the models with published tokenizers are not the same list, and pairing them would be a guess.

From the directory
English$5.74a month · $69 a year5.7M tokens a month
العربية$6.80a month · $82 a year6.8M tokens a month+19% against English
中文$22a month · $266 a year22M tokens a month+287% against English

Input tokens only, at 1.15 tokens per English word measured on this corpus. Output is priced separately and usually higher; the same multiplier applies to it.

Not measured

Models that publish no tokenizer

A count can only be shown where the tokenizer is published. These are left out rather than estimated:

  • ClaudeAnthropic publishes no tokenizer, in any form.
  • GeminiGoogle publishes no Gemini tokenizer; counts come only from a metered API.
  • ALLaMSaudi Arabia's national model publishes no tokenizer file at its address.
  • Aya ExpanseGated. Aya 101, measured above, is a different and older model.
Provenance

Where each vocabulary came from

Meta, Google and Inception gate their own repositories, so three of the rows above were read from republications rather than from the owner. A copy is only as good as its faithfulness, so each was checked against a second republisher with no connection to the first: same vocabulary size, and the same token identifiers for the same text, down to the number. Where two unrelated parties agree exactly, the copy is the original.

  • ALLaM 7Bread from JasperV13/Yehia-7B-DPO-Reasoning-previewchecked against ALLaM-AI/ALLaM-7B-Instruct-preview tokenizer.model, byte-identical
  • Gemma 2 9Bread from unsloth/gemma-2-9b-itchecked against rinna/gemma-2-baku-2b
  • Gemma 2 2Bread from unsloth/gemma-2-2b-itchecked against rinna/gemma-2-baku-2b
  • Llama 3.xread from NousResearch/Meta-Llama-3.1-8B-Instructchecked against unsloth/Llama-3.3-70B-Instruct
How this was measured

Method

Comparing tokenizers across languages requires parallel text carrying broadly the same meaning. This benchmark uses the same 6 articles of the Universal Declaration of Human Rights from the United Nations' English, Arabic and Chinese pages because they are stable, substantial and closely aligned. The subject is incidental: this is a tokenization benchmark, not an assessment of human-rights knowledge or policy.

The sentence shown split token by token is deliberately separate from the benchmark corpus. It is quoted from the President's site, which publishes it in Arabic and English, and is included only to make token splitting visible. It is unusually compact in Arabic, so using it for the ratios would bias the comparison. Every ratio and cost figure here comes from the declaration; the quote supplies only the split you can see.

Each available published tokenizer implementation is run over that text: OpenAI's rank files for the GPT tokenizers and tokenizer files from Hugging Face for the rest. Where an owner's repository is gated, the provenance section identifies the documented, independently checked republication used instead. The token counts are exact for this corpus; no rule of thumb is used.

Ratios are computed over the whole corpus rather than averaged across articles, so a long article counts for more than a short one — which is how a bill works.

The register is formal and legal. Ratios move somewhat with register, so treat these as the shape of the problem rather than a figure to quote to two decimals on your own text.

Corpus Universal Declaration of Human Rights · Measured Sep 28, 2026