OpenAI

GPT-6 Sol

8.8/10
Maker
OpenAI
Origin
United States
Released
Sep 22, 2026

Strengths & weaknesses

  • 5th of 133 on LMArena's WebDev arena, OpenAI's strongest showing for building working web apps after Astra
  • 93.5% factual consistency on Vectara's hallucination leaderboard, ahead of GPT-6 Astra's 91.3%
  • 156th of 409 on LMArena's overall text arena — in blind comparisons people prefer a great many cheaper models
  • Sits below Astra in OpenAI's own line-up, so it doesn't carry the flagship's headline capabilities

Evaluation

100% of weight measured

Scored on GPT-6 Sol (Max) · sources as of Sep 26, 2026

Confidence band 8.8–9.0 · a model inside this range isn’t meaningfully apart from this one

Price & availability

Measured the same way, deliberately kept out of the score: what a model costs doesn’t change what it can do.

w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works

More in Language Models

View all →
Meta
  • Meta's first frontier line since it moved off Llama, natively multimodal across text, image and video at $1.25/$4.25 per million tokens
  • Meta reports roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on the same agentic work
  • Proprietary, ending the open-weight position Llama gave Meta, and served by a single provider on OpenRouter
  • 10th of 409 on LMArena's overall text arena — strong, but behind the Anthropic and Google models above it
Google
  • Highest-rated model in the directory on LMArena's hard prompts after Anthropic's Opus line, on 14,320 votes
  • Google reports 90.8% on Terminal-Bench 2.1, up from 81.6% for 3.7 Flash, at $0.75/$3.75 per million tokens
  • Built on 3.7 Flash rather than a new base model, and Google tells you to stay on 3.7 for efficiency-first work
  • Spends more thinking tokens by design, so the headline price understates what a task actually costs
Moonshot AI
9.4/10
  • The largest open-weight release yet at 2.8T parameters with 104B active, served by 17 independent providers
  • Moonshot reports 81.2 on FrontierSWE and 88.3 on Terminal-Bench 2.1, and LMArena ranks it 8th of 133 on WebDev
  • $3/$15 per million tokens, several times the price of the open-weight models it competes with
  • Its licence is only open to a point: model-as-a-service businesses over $20M a year need a separate agreement
Alibaba
  • 6th of 133 on LMArena's WebDev arena, the highest-rated non-frontier model for building working web apps
  • 2.4T mixture-of-experts with about 95B active and a 1M-token context, at $2/$6 per million tokens
  • Loses to GLM-5.3 on seven of the eight benchmarks both makers publish, at several times the price
  • Served by a single provider on OpenRouter, so there is no fallback if it is unavailable