Percent scores map straight to marks
When a public benchmark reports a percentage, the mark is that percentage divided by ten. 72% becomes 7.2. The mark doesn’t move when other models are added.
SWE-bench Verified: 75.6% → 7.6
A score is only worth something if you can see how it was made. This page explains every step, first in plain words, then in full detail.
Methodology v1.2 · Sep 19, 2026 · 32 of 100 models scored. The rest stay blank until there’s enough evidence.
You wouldn’t grade a fish on climbing trees. Image models are graded on things like picture quality and following instructions; chat models on reasoning, accuracy, and speed. Each category has its own subjects.
Reasoning: 9. Writing: 7. Cost: 5. Every mark comes from a public measurement or a published checklist, never from a gut feeling.
For a chat model, reasoning counts for 30% of the final grade and writing for 10%. The weight of every subject is shown next to it (w30, w10).
Strong marks in heavy subjects pull the score up more than strong marks in light ones. That’s the number in the corner of a card, like 8.2/10.
A model needs marks covering at least 80% of what counts before it gets a final score, so one lucky subject can’t make it look amazing.
Companies don’t pay to be listed or scored, and they don’t see scores before they’re published. If a number is wrong, the fix is better evidence, not a phone call.
Every criterion is scored with one of three rules. The rule used, the source, and the date appear next to each mark on the model’s page, so any number can be checked.
When a public benchmark reports a percentage, the mark is that percentage divided by ten. 72% becomes 7.2. The mark doesn’t move when other models are added.
SWE-bench Verified: 75.6% → 7.6
Arena ratings, price, and context length have no natural top score, so a model is ranked against the directory’s other measured models in its category. Its mark is 10 × the share of those models (counting itself) that it matches or beats, so the best gets 10. Cheaper counts as better.
Matches or beats 8 of 10 measured models → 8.0
Things like licensing or data controls can’t be benchmarked, so each gets a five-item checklist worth two points per item. Each checklist is published on this page before any model is marked against it.
4 of 5 checklist items met → 8.0
All independent, all public. Their data is copied into a dated snapshot, so scores only change when the snapshot is refreshed, and the date shows on every mark.
Human-preference ratings from blind, side-by-side votes: text (hard prompts, creative writing), web development, text-to-image, image editing, text-to-video, and image-to-video.
Arenas published 2026-09-04 to 2026-09-15
How often a model sticks to the facts in a document it summarises (factual consistency rate).
Share of real GitHub issues resolved. Only runs using the same open harness (mini-SWE-agent) are counted, so results are comparable.
List prices, context windows, supported features, and the number of independent providers serving each model.
Two points per item, checked against OpenRouter’s data.
Here is an example language model. It isn’t a real model; the marks are made up to show the calculation. The numbers below are computed by the same code that scores the directory.
| Criterion | Mark | Weight | Mark × weight |
|---|---|---|---|
| Reasoning | 9 | 30 | 270 |
| Accuracy | 8 | 20 | 160 |
| Writing | 7 | 10 | 70 |
| Long context | 6 | 10 | 60 |
| Cost efficiency | 5 | 20 | 100 |
| Deployability | 8 | 10 | 80 |
| Total | 100 | 740 |
740 ÷ 100 = 7.4
Add up each mark multiplied by its weight, then divide by the total weight. Rounded to one decimal place, only at the end.
With only Reasoning and Writing marked, the evidence covers 40% of the weight, below the 80% bar. So there’s no overall score, and the card says “in assessment” instead. Once a model clears the bar, the average uses only the marks that exist and their weights.
They rate how strong the proof is that a use case pays off, not how good a model is. See the studies
How much each use case is used, from the Anthropic Economic Index. Popular isn’t the same as good, so usage never counts towards a score.
See the top 100 use casesWeights are relative within a rubric and shown as a share of the total. They are fixed before scoring; if one changes, every affected score is recalculated and the change is dated.
| Criterion · measured by | Weight |
|---|---|
Reasoning Preferred by people on hard, multi-step prompts. Rule BLMArena text arena, hard prompts | 30% |
Accuracy Sticks to the facts it was given, with few hallucinations. Rule AVectara hallucination leaderboard, factual consistency rate | 20% |
Writing Preferred by people for creative and long-form writing. Rule BLMArena text arena, creative writing | 10% |
Long context How much text it can take in at once. Rule BOpenRouter, context window | 10% |
Cost efficiency Price per token; cheaper scores higher. Rule BOpenRouter, blended price per token | 20% |
Deployability Open weights, choice of hosts, and integration features. Rule CDeployability checklist (OpenRouter data) | 10% |
| Criterion · measured by | Weight |
|---|---|
Code quality Share of real GitHub issues it resolves on its own. Rule ASWE-bench Verified, % resolved (mini-SWE-agent harness) | 35% |
Web development Preferred by people for building working web apps. Rule BLMArena WebDev arena | 35% |
Cost efficiency Price per token; cheaper scores higher. Rule BOpenRouter, blended price per token | 15% |
Codebase context How much code it can take in at once. Rule BOpenRouter, context window | 15% |
| Criterion · measured by | Weight |
|---|---|
Image quality Preferred by people across all kinds of image prompts. Rule BLMArena text-to-image arena | 40% |
Text rendering Preferred by people for images that contain text. Rule BLMArena text-to-image, text rendering | 20% |
Commercial design Preferred by people for ads, product shots, and branding. Rule BLMArena text-to-image, commercial design | 20% |
Editing Preferred by people for editing existing images. Rule BLMArena image-edit arena | 20% |
| Criterion · measured by | Weight |
|---|---|
Text-to-video Preferred by people for video made from a written prompt. Rule BLMArena text-to-video arena | 60% |
Image-to-video Preferred by people for video made from a starting image. Rule BLMArena image-to-video arena | 40% |
| Criterion · measured by | Weight |
|---|---|
Output quality Naturalness, fidelity, and musicality — or transcription accuracy. Not scored yet: no independent source | 35% |
Control Steering of voice, emotion, structure, or style. Not scored yet: no independent source | 20% |
Language coverage Breadth and quality across languages and accents. Not scored yet: no independent source | 15% |
Latency Fast enough for real-time or high-volume use. Not scored yet: no independent source | 10% |
Licensing & safety Commercial rights and consent safeguards. Not scored yet: no independent source | 10% |
Cost efficiency Price per minute or per track. Not scored yet: no independent source | 10% |
| Criterion · measured by | Weight |
|---|---|
Answer grounding Accurate answers anchored in retrieved sources. Not scored yet: no independent source | 30% |
Citations Clear, checkable references back to sources. Not scored yet: no independent source | 20% |
Integrations Connectors to company data, clouds, and tools. Not scored yet: no independent source | 20% |
Governance Permissions, audit, data residency, and admin controls. Not scored yet: no independent source | 15% |
Cost efficiency Predictable pricing at organisational scale. Not scored yet: no independent source | 15% |
| Criterion · measured by | Weight |
|---|---|
Functional depth How well it covers the real workflows of its function. Not scored yet: no independent source | 30% |
Integration Fits existing systems of record and data. Not scored yet: no independent source | 20% |
Time to value Speed and effort of rollout and adoption. Not scored yet: no independent source | 20% |
Scalability Holds up across headcount, regions, and volume. Not scored yet: no independent source | 15% |
Total cost Licence plus implementation and upkeep. Not scored yet: no independent source | 15% |
If a rule or weight changes, it’s listed here with its date, and every affected score is recalculated.