Scoring data updated
Next check in--:--:--
Beta · still being built

100 AI models,
measured, not estimated.

A field guide to the models shaping AI across language, code, image, video, audio, and the enterprise. It covers what each one genuinely leads at, where it still falls short, where it comes from, and when it shipped.

Models tracked
100
Categories
06
Countries of origin
12
Use cases
1000
Global AI adoption

Where generative AI is used most

Two independent studies see different leaders because they measure different behavior. Microsoft puts the UAE first for estimated generative-AI use; Cybernews puts it third for mobile AI-app downloads. The percentages are kept separate rather than averaged.

Microsoft AI diffusion

AI user share · June 2026

  1. 01United Arab Emirates73.3%+3.2pp
  2. 02Singapore64.3%+0.9pp
  3. 03Ireland49.9%+1.5pp
  4. 04France49.6%+1.8pp
  5. 05Norway49.4%+0.8pp
  6. 06Spain45.1%+0.9pp
  7. 07New Zealand44.1%+1.1pp
  8. 08United Kingdom43.7%+1.5pp
  9. 09Netherlands43.5%+1.4pp
  10. 10Qatar43.5%+1.7pp

Worldwide 18.8%

Based on aggregated, anonymized Microsoft telemetry, adjusted for operating-system and device share, internet penetration, and population. This measures generative-AI use among people aged 15–64; it is not a measure of every form of AI adoption.

Q2 2026 reportTechnical method

Cybernews app adoption

AI-app downloads per population · 2025

  1. 01Singapore66%
  2. 02Chile60%
  3. 03United Arab Emirates56%

Cybernews divides downloads of the 100 most popular AI-first apps from Google Play and Apple’s App Store by total population. Downloads are not unique active users, and browser, desktop and sideloaded tools are excluded.

Cybernews studyTechnical method

These rankings are not combined: their populations, periods and observed behavior differ.

Leaderboard

Highest scored

40 models have enough independent measurement to be scored. Each category is marked on its own rubric, so this mixed view is a rough guide — use the tabs to compare like with like.

Highest scored

  1. 01Claude Opus 5Anthropic · United States · Language ModelsDriven by reasoning — 1,526 rating on LMArena’s hard prompts, joint highest among language models.9.7
  2. 02Claude Opus 5.5Anthropic · United States · Language ModelsDriven by reasoning — 1,536 rating on LMArena’s hard prompts, joint highest among language models.Statistically tied with #19.7
  3. 03Claude Fable 5.1Anthropic · United States · Language ModelsDriven by reasoning — 1,526 rating on LMArena’s hard prompts, third among language models.Statistically tied with #19.6
  4. 04Claude Fable 5Anthropic · United States · Language ModelsDriven by reasoning — 1,508 rating on LMArena’s hard prompts, fourth among language models.9.5
  5. 05Gemini 3.8 FlashGoogle · United States · Language ModelsDriven by writing — 1,489 rating on LMArena’s creative writing, fifth among language models.Statistically tied with #49.5
  6. 06Kimi K3Moonshot AI · China · Language ModelsDriven by reasoning — 1,496 rating on LMArena’s hard prompts, 8th of 25 language models.Statistically tied with #49.4
  7. 07MiMo v2.6Xiaomi · China · Language ModelsDriven by reasoning — 1,512 rating on LMArena’s hard prompts, joint fifth among language models.Statistically tied with #49.4
  8. 08Muse Spark 1.3Meta · United States · Language ModelsDriven by reasoning — 1,504 rating on LMArena’s hard prompts, joint fifth among language models.Statistically tied with #49.4
  9. 09Qwen3.8 MaxAlibaba · China · Language ModelsDriven by reasoning — 1,495 rating on LMArena’s hard prompts, 8th of 25 language models.Statistically tied with #49.4
  10. 10GLM-5.3Zhipu AI · China · Language ModelsDriven by reasoning — 1,489 rating on LMArena’s hard prompts, 10th of 25 language models.Statistically tied with #49.3

Only the top is published. Further down a ranking the marks are mostly older models kept for reference, or models with just enough measurement to be scored, so a “lowest scored” table would say more about the gaps in the evidence than about the products. Each model’s own page shows its marks and what’s still missing.

The score measures capability only: price and availability are measured too, but shown on each model’s page rather than folded into this number. Marks use the low end of each source’s confidence interval, and models whose intervals overlap are marked as tied, because the evidence can’t separate them. Models with too little independent measurement are under evaluation and don’t appear here at all. How scoring works

Arabic leaderboard

Models built in the Arab world

Ranked on how well they handle Arabic — which the scores above say nothing about, because the benchmarks behind them are run in English. Two independent leaderboards publish an Arabic measurement in a form that can be read and checked. Both are shown for every model, and both are combined into the figure the table is ordered by.

AraGen v3 · 67 · Apr 4, 2026OALL v2 · 163 · Mar 2, 2026

The large numbers are each board’s own result, on its own scale. The table is ordered by the combined figure beside them, which is why a model can show a high number here and still sit low: the combined figure carries the cost of converting between the two scales, and the row says which board measured it.

  1. 01
    Jais 2 70B ChatMBZUAICerebrasInception · UAEInception's flagship, trained with MBZUAI and Cerebras on Arabic and English, with a vocabulary built for Modern Standard Arabic, regional dialects and Arabic–English code-switching.
    • The highest-placed Arab-world model on AraGen v3: 18th of 67 at 0.52 3C3H, ahead of Llama 3.3 70B and Mistral Large on Arabic free-form answers
    • One of the few Arab-world models trained from scratch rather than adapted from a Western base, and open-weight under Apache 2.0
    • All 17 models above it on AraGen are non-Arab frontier systems, and GPT-5 leads it by 0.32 on the same scale
    • 8,192-token context, and OALL has never run it — so there is no native-benchmark reading to check the judged one against
    AraGen0.52 3C3H (0.48–0.57) · #18 of 67 · Dec 14, 2025OALLnot measured
    AraGen4.8OALL—not measuredCombined4.8
  2. 02
    Falcon H1 34B InstructTII · UAEAbu Dhabi's hybrid Transformer–Mamba model from TII, at 34B parameters.
    • Open-weight and general-purpose across 18 languages, so one model serves Arabic and the rest rather than sitting beside a separate Arabic build
    • 28th of 67 on AraGen v3 at 0.46 3C3H, the second-best of the Arab-world models listed here
    • Below every frontier model on AraGen, and 0.07 behind Jais 2 70B on the same 3C3H scale
    • The Arabic-specialised Falcon-H1-Arabic that TII has since released is on neither board, so what is measured here is the general model
    AraGen0.46 3C3H (0.41–0.51) · #28 of 67 · Sep 28, 2025OALLnot measuredStatistically tied with #1
    AraGen4.1OALL—not measuredCombined4.1
  3. 03
    Karnakconverted from OALLApplied Innovation Center · EgyptEgypt's state model, built at the Applied Innovation Center and tuned on material published in Egypt and across the Arab world, covering Modern Standard and Egyptian colloquial Arabic.
    • Top of OALL v2: 1st of 163 at 79.3% across the seven Arabic benchmarks, ahead of every Arabic fine-tune measured there
    • The Applied Innovation Center's other model, AIC-1, sits 10th on the same board: two of OALL's top ten come from one Egyptian lab
    • AraGen has never run it, so its free-form Arabic has no judged measurement and its combined figure rests entirely on the conversion
    • A tune of an open-weight base rather than a model trained from scratch, and Karnak Chat is still a public pilot
    AraGennot measuredOALL79.3% average · #1 of 163 · Feb 15, 2026Statistically tied with #1
    AraGen—not measuredOALL7.9Combined2.2
  4. 04
    Fanar 1 9B InstructQCRI · QatarQatar's model, continually pretrained from Gemma 2 on a trillion Arabic and English tokens at QCRI, covering Gulf, Levantine and Egyptian dialects and deliberately aligned to Islamic values and Arab culture.
    • The only model here both boards have measured, so its combined figure needs no conversion between scales
    • 25th of 163 on OALL v2 at 70.3% — ahead of AceGPT-v2 70B and Qwen2 72B at a fraction of their size
    • 54th of 67 on AraGen v3 at 0.16 3C3H: it answers native benchmark questions well and free-form judged ones poorly
    • 4,096-token context, and the 27B Fanar 2.0 QCRI has since announced is on neither board
    AraGen0.16 3C3H (0.12–0.20) · #54 of 67 · Sep 30, 2025OALL70.3% average · #25 of 163 · Jun 7, 2025
    AraGen1.2OALL7.0Combined1.3
  5. 05
    Yehia 7Bconverted from OALLNavid AI · Saudi ArabiaNavid AI's fine-tune of ALLaM, Saudi Arabia's national model — a startup's attempt to improve on a state lab's release.
    • 48th of 163 on OALL v2 at 65.7%, edging the ALLaM base it was fine-tuned from, which sits 53rd at 65.3%
    • Trained with GRPO against the same six qualities AraGen judges, aiming at free-form answer quality rather than benchmark accuracy
    • AraGen has never run it, so the quality it was explicitly trained for has never been measured on it
    • A 7B fine-tune of another maker's model, and two other Yehia variants sit within 0.2 points of it on OALL — inside the noise
    AraGennot measuredOALL65.7% average · #48 of 163 · Mar 3, 2025Statistically tied with #4
    AraGen—not measuredOALL6.6Combined1.3
  6. 06
    ALLaM 7B Instructconverted from OALLSDAIA · Saudi ArabiaSaudi Arabia's national model, trained from scratch at SDAIA's National Center for AI, and the base Yehia 7B and its variants are built on.
    • Trained from scratch at SDAIA on 5.2 trillion tokens, 1.2 trillion of them Arabic and English, and open-weight under Apache 2.0
    • 53rd of 163 on OALL v2 at 65.3%, ahead of Inception's 30B jais-family chat model at four times the size
    • 4,096-token context, and still published as a preview
    • AraGen has never run it, and on the one board that has, a community fine-tune of it — Yehia 7B — scores higher
    AraGennot measuredOALL65.3% average · #53 of 163 · Feb 19, 2025Statistically tied with #4
    AraGen—not measuredOALL6.5Combined1.2

The two scales aren’t interchangeable, so they’re not simply averaged. A line is fitted across the 25 models both boards have run, where the agreement is moderate rather than tight (r = 0.67). Converting an OALL result into AraGen’s terms costs ±2.2 marks at 95%, and the two figures are then weighted by how firm each one is, so a directly measured score counts far more than a converted one.

A model only one board has measured therefore carries a wide band, is marked at the low end of it like every other mark on the site, and ties with much of the table. That is the honest reading rather than a fault in the table: its Arabic is as good as the board that measured it says, and what that means on the other scale isn’t known. Read the two source columns, not just the combined figure — where they disagree, the disagreement is the finding.

Each board’s own figure follows the same rule as the rest of the site: a published percentage divided by 10, never an estimate. These models are kept apart from the directory’s 100 because they are judged on different benchmarks, against a different field — a place here says nothing about where a model stands globally, and a place there says nothing about its Arabic. How scoring works

Multilingual leaderboard

Outside English

LMArena splits the same blind, side-by-side votes by the language of the prompt. This is the directory’s models on the prompts that weren’t in English — the same voters, the same models, different work. No arena breaks Arabic out on its own, which is why the board above is built from Arabic benchmarks instead.

LMArena406 models measured, published Oct 2, 2026

  1. 01Claude Fable 5.1Anthropic · United States#2 of 406 outside English · #7 of 412 in English1,507 rating (1,500–1,515) · 7,005 votes9.4
  2. 02Claude Opus 5Anthropic · United States#3 of 406 outside English · #5 of 412 in English1,501 rating (1,495–1,507) · 17,233 votesStatistically tied with #19.3
  3. 03Claude Opus 5.5Anthropic · United States#4 of 406 outside English · #4 of 412 in English1,503 rating (1,492–1,514) · 2,829 votesStatistically tied with #19.3
  4. 04Gemini 3.8 FlashGoogle · United States#7 of 406 outside English · #10 of 412 in English1,488 rating (1,482–1,494) · 15,930 votes9.2
  5. 05Claude Fable 5Anthropic · United States#10 of 406 outside English · #8 of 412 in English1,482 rating (1,477–1,487) · 21,907 votesStatistically tied with #49.1
  6. 06Muse Spark 1.3Meta · United States#12 of 406 outside English · #13 of 412 in English1,482 rating (1,475–1,490) · 7,408 votesStatistically tied with #49.1
  7. 07Gemini 3 ProGoogle · United States#16 of 406 outside English · #25 of 412 in English1,474 rating (1,470–1,479) · 25,020 votes9.0
  8. 08Kimi K3Moonshot AI · China#26 of 406 outside English · #17 of 412 in English1,465 rating (1,460–1,471) · 17,482 votesStatistically tied with #78.9
  9. 09MiMo v2.6Xiaomi · China#22 of 406 outside English · #11 of 412 in English1,475 rating (1,463–1,486) · 2,566 votesStatistically tied with #78.9
  10. 10Qwen3.8 MaxAlibaba · China#19 of 406 outside English · #12 of 412 in English1,471 rating (1,465–1,477) · 14,348 votesStatistically tied with #78.9

28 of the directory’s models clear the vote floor on this board. The top 10 are listed, highest mark first.

The two places are counted in the same field of models, so they can be read against each other — but read them gently. Around any given place the ratings sit inside one another’s confidence bands, so a move of even twenty places usually can’t be told apart from noise. A model far from its English place is worth a look; a few places either way is not.

Marked on the arena’s own fixed scale at the low end of each published interval, the same rule the directory’s scores follow, and a rating built on fewer than 500 votes is treated as unmeasured rather than weak. How scoring works

The context window, every model in view

Index

In this directory

Changelog

Latest releases

  1. Claude Sonnet 5.5AnthropicLanguage Models
  2. Claude Opus 5.5AnthropicLanguage Models
  3. GPT-6 SolOpenAILanguage Models
  4. GPT-6 LunaOpenAILanguage Models
  5. MiMo v2.6XiaomiLanguage Models
  6. Grok 4.7SpaceXAILanguage Models
  7. Step 5 PreviewStepFunCoding & Agentic
  8. DeepSeek V4.1 FlashDeepSeekLanguage Models
Provenance

Where the models come from

Each maker’s headquarters country. The frontier is concentrated in two places, but Europe, the Gulf, Israel, and Asia-Pacific each ship models worth knowing.

  • United States56
  • China23
  • United Kingdom5
  • Israel4
  • Australia2
  • Canada2
  • France2
  • Germany1
  • Luxembourg1
  • Russia1
  • South Korea1
  • Spain1