SpaceXAI

Grok 4.7

8.2/10
Maker
SpaceXAI
Origin
United States
Released
Sep 21, 2026

Strengths & weaknesses

  • SpaceXAI reports Terminal-Bench 4.0 nearly doubling to 38.0% from Grok 4.6's 20.3% on multi-hour terminal tasks
  • Holds Grok 4.6's $2/$6 per million tokens despite a larger base model and a longer reinforcement-learning run
  • 151st of 399 on LMArena's hard-prompts arena — the gains its maker reports on agentic benchmarks don't show up in blind human preference
  • SpaceXAI's own figures put it behind Claude Fable 5.1 on CursorBench 4.0, 46.3% against 51.8%

Evaluation

80% of weight measured

Scored on Grok 4.7 (xHigh) · sources as of Sep 26, 2026

Confidence band 8.2–8.5 · a model inside this range isn’t meaningfully apart from this one

Price & availability

Measured the same way, deliberately kept out of the score: what a model costs doesn’t change what it can do.

w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works

More in Language Models

View all →
Meta
  • Meta's first frontier line since it moved off Llama, natively multimodal across text, image and video at $1.25/$4.25 per million tokens
  • Meta reports roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on the same agentic work
  • Proprietary, ending the open-weight position Llama gave Meta, and served by a single provider on OpenRouter
  • 10th of 409 on LMArena's overall text arena — strong, but behind the Anthropic and Google models above it
Google
  • Highest-rated model in the directory on LMArena's hard prompts after Anthropic's Opus line, on 14,320 votes
  • Google reports 90.8% on Terminal-Bench 2.1, up from 81.6% for 3.7 Flash, at $0.75/$3.75 per million tokens
  • Built on 3.7 Flash rather than a new base model, and Google tells you to stay on 3.7 for efficiency-first work
  • Spends more thinking tokens by design, so the headline price understates what a task actually costs
Moonshot AI
9.4/10
  • The largest open-weight release yet at 2.8T parameters with 104B active, served by 17 independent providers
  • Moonshot reports 81.2 on FrontierSWE and 88.3 on Terminal-Bench 2.1, and LMArena ranks it 8th of 133 on WebDev
  • $3/$15 per million tokens, several times the price of the open-weight models it competes with
  • Its licence is only open to a point: model-as-a-service businesses over $20M a year need a separate agreement
Alibaba
  • 6th of 133 on LMArena's WebDev arena, the highest-rated non-frontier model for building working web apps
  • 2.4T mixture-of-experts with about 95B active and a 1M-token context, at $2/$6 per million tokens
  • Loses to GLM-5.3 on seven of the eight benchmarks both makers publish, at several times the price
  • Served by a single provider on OpenRouter, so there is no fallback if it is unavailable