100 AI models,100 AI models, measured, not estimated.
A field guide to the models shaping AI across language, code, image, video, audio, and the enterprise. It covers what each one genuinely leads at, where it still falls short, where it comes from, and when it shipped.
Where generative AI is used mostWhere generative AI is used most
Two independent studies see different leaders because they measure different behavior. Microsoft puts the UAE first for estimated generative-AI use; Cybernews puts it third for mobile AI-app downloads. The percentages are kept separate rather than averaged.
Microsoft AI diffusion
AI user share · June 2026
01United Arab Emirates73.3%+3.2pp
02Singapore64.3%+0.9pp
03Ireland49.9%+1.5pp
04France49.6%+1.8pp
05Norway49.4%+0.8pp
06Spain45.1%+0.9pp
07New Zealand44.1%+1.1pp
08United Kingdom43.7%+1.5pp
09Netherlands43.5%+1.4pp
10Qatar43.5%+1.7pp
Cybernews app adoption
AI-app downloads per population · 2025
01Singapore66%
02Chile60%
03United Arab Emirates56%
These rankings are not combined: their populations, periods and observed behavior differ.
Leaderboard
Highest scoredHighest scored
40 models have enough independent measurement to be scored. Each category is marked on its own rubric, so this mixed view is a rough guide — use the tabs to compare like with like.
Only the top is published. Further down a ranking the marks are mostly older models kept for reference, or models with just enough measurement to be scored, so a “lowest scored” table would say more about the gaps in the evidence than about the products. Each model’s own page shows its marks and what’s still missing.
The score measures capability only: price and availability are measured too, but shown on each model’s page rather than folded into this number. Marks use the low end of each source’s confidence interval, and models whose intervals overlap are marked as tied, because the evidence can’t separate them. Models with too little independent measurement are under evaluation and don’t appear here at all. How scoring works
Arabic leaderboard
Models built in the Arab worldModels built in the Arab world
Ranked on how well they handle Arabic — which the scores above say nothing about, because the benchmarks behind them are run in English. Two independent leaderboards publish an Arabic measurement in a form that can be read and checked. Both are shown for every model, and both are combined into the figure the table is ordered by.
The large numbers are each board’s own result, on its own scale. The table is ordered by the combined figure beside them, which is why a model can show a high number here and still sit low: the combined figure carries the cost of converting between the two scales, and the row says which board measured it.
01
Jais 2 70B Chat↗MBZUAICerebrasInception · UAEInception's flagship, trained with MBZUAI and Cerebras on Arabic and English, with a vocabulary built for Modern Standard Arabic, regional dialects and Arabic–English code-switching.
+The highest-placed Arab-world model on AraGen v3: 18th of 67 at 0.52 3C3H, ahead of Llama 3.3 70B and Mistral Large on Arabic free-form answers
+One of the few Arab-world models trained from scratch rather than adapted from a Western base, and open-weight under Apache 2.0
−All 17 models above it on AraGen are non-Arab frontier systems, and GPT-5 leads it by 0.32 on the same scale
−8,192-token context, and OALL has never run it — so there is no native-benchmark reading to check the judged one against
AraGen0.52 3C3H (0.48–0.57) · #18 of 67 · Dec 14, 2025OALLnot measured
AraGen4.8OALL—not measuredCombined4.8
02
Falcon H1 34B Instruct↗TII · UAEAbu Dhabi's hybrid Transformer–Mamba model from TII, at 34B parameters.
+Open-weight and general-purpose across 18 languages, so one model serves Arabic and the rest rather than sitting beside a separate Arabic build
+28th of 67 on AraGen v3 at 0.46 3C3H, the second-best of the Arab-world models listed here
−Below every frontier model on AraGen, and 0.07 behind Jais 2 70B on the same 3C3H scale
−The Arabic-specialised Falcon-H1-Arabic that TII has since released is on neither board, so what is measured here is the general model
AraGen0.46 3C3H (0.41–0.51) · #28 of 67 · Sep 28, 2025OALLnot measuredStatistically tied with #1
AraGen4.1OALL—not measuredCombined4.1
03
Karnak↗converted from OALLApplied Innovation Center · EgyptEgypt's state model, built at the Applied Innovation Center and tuned on material published in Egypt and across the Arab world, covering Modern Standard and Egyptian colloquial Arabic.
+Top of OALL v2: 1st of 163 at 79.3% across the seven Arabic benchmarks, ahead of every Arabic fine-tune measured there
+The Applied Innovation Center's other model, AIC-1, sits 10th on the same board: two of OALL's top ten come from one Egyptian lab
−AraGen has never run it, so its free-form Arabic has no judged measurement and its combined figure rests entirely on the conversion
−A tune of an open-weight base rather than a model trained from scratch, and Karnak Chat is still a public pilot
AraGennot measuredOALL79.3% average · #1 of 163 · Feb 15, 2026Statistically tied with #1
AraGen—not measuredOALL7.9Combined2.2
04
Fanar 1 9B Instruct↗QCRI · QatarQatar's model, continually pretrained from Gemma 2 on a trillion Arabic and English tokens at QCRI, covering Gulf, Levantine and Egyptian dialects and deliberately aligned to Islamic values and Arab culture.
+The only model here both boards have measured, so its combined figure needs no conversion between scales
+25th of 163 on OALL v2 at 70.3% — ahead of AceGPT-v2 70B and Qwen2 72B at a fraction of their size
−54th of 67 on AraGen v3 at 0.16 3C3H: it answers native benchmark questions well and free-form judged ones poorly
−4,096-token context, and the 27B Fanar 2.0 QCRI has since announced is on neither board
AraGen0.16 3C3H (0.12–0.20) · #54 of 67 · Sep 30, 2025OALL70.3% average · #25 of 163 · Jun 7, 2025
AraGen1.2OALL7.0Combined1.3
05
Yehia 7B↗converted from OALLNavid AI · Saudi ArabiaNavid AI's fine-tune of ALLaM, Saudi Arabia's national model — a startup's attempt to improve on a state lab's release.
+48th of 163 on OALL v2 at 65.7%, edging the ALLaM base it was fine-tuned from, which sits 53rd at 65.3%
+Trained with GRPO against the same six qualities AraGen judges, aiming at free-form answer quality rather than benchmark accuracy
−AraGen has never run it, so the quality it was explicitly trained for has never been measured on it
−A 7B fine-tune of another maker's model, and two other Yehia variants sit within 0.2 points of it on OALL — inside the noise
AraGennot measuredOALL65.7% average · #48 of 163 · Mar 3, 2025Statistically tied with #4
AraGen—not measuredOALL6.6Combined1.3
06
ALLaM 7B Instruct↗converted from OALLSDAIA · Saudi ArabiaSaudi Arabia's national model, trained from scratch at SDAIA's National Center for AI, and the base Yehia 7B and its variants are built on.
+Trained from scratch at SDAIA on 5.2 trillion tokens, 1.2 trillion of them Arabic and English, and open-weight under Apache 2.0
+53rd of 163 on OALL v2 at 65.3%, ahead of Inception's 30B jais-family chat model at four times the size
−4,096-token context, and still published as a preview
−AraGen has never run it, and on the one board that has, a community fine-tune of it — Yehia 7B — scores higher
AraGennot measuredOALL65.3% average · #53 of 163 · Feb 19, 2025Statistically tied with #4
AraGen—not measuredOALL6.5Combined1.2
The two scales aren’t interchangeable, so they’re not simply averaged. A line is fitted across the 25 models both boards have run, where the agreement is moderate rather than tight (r = 0.67). Converting an OALL result into AraGen’s terms costs ±2.2 marks at 95%, and the two figures are then weighted by how firm each one is, so a directly measured score counts far more than a converted one.
A model only one board has measured therefore carries a wide band, is marked at the low end of it like every other mark on the site, and ties with much of the table. That is the honest reading rather than a fault in the table: its Arabic is as good as the board that measured it says, and what that means on the other scale isn’t known. Read the two source columns, not just the combined figure — where they disagree, the disagreement is the finding.
Each board’s own figure follows the same rule as the rest of the site: a published percentage divided by 10, never an estimate. These models are kept apart from the directory’s 100 because they are judged on different benchmarks, against a different field — a place here says nothing about where a model stands globally, and a place there says nothing about its Arabic. How scoring works
Multilingual leaderboard
Outside EnglishOutside English
LMArena splits the same blind, side-by-side votes by the language of the prompt. This is the directory’s models on the prompts that weren’t in English — the same voters, the same models, different work. No arena breaks Arabic out on its own, which is why the board above is built from Arabic benchmarks instead.
LMArena·406 models measured, published Oct 2, 2026
28 of the directory’s models clear the vote floor on this board. The top 10 are listed, highest mark first.
The two places are counted in the same field of models, so they can be read against each other — but read them gently. Around any given place the ratings sit inside one another’s confidence bands, so a move of even twenty places usually can’t be told apart from noise. A model far from its English place is worth a look; a few places either way is not.
Marked on the arena’s own fixed scale at the low end of each published interval, the same rule the directory’s scores follow, and a rating built on fewer than 500 votes is treated as unmeasured rather than weak. How scoring works
Where the models come fromWhere the models come from
Each maker’s headquarters country. The frontier is concentrated in two places, but Europe, the Gulf, Israel, and Asia-Pacific each ship models worth knowing.