StepFun

Step 5 Preview

Under evaluation
Maker
StepFun
Origin
China
Released
Sep 20, 2026

Strengths & weaknesses

  • 600B mixture-of-experts with 27B active and a 1M-token context, built for long-horizon agentic work
  • $1.00 input and $2.70 output per million tokens, with open weights promised for October 2026
  • Generated roughly 1.7× the median output tokens on Artificial Analysis's index, so real cost runs well above the headline rate
  • A preview with no independent coding-benchmark result yet, so it carries no overall score here

Evaluation

In assessment · 40% measured

Scored on Step 5 Preview (High) · sources as of Sep 26, 2026

  • Code qualityw45—

    No independent measurement yet

  • Web developmentw406.6

    Rule B1,565 rating (95% 1,548–1,583) from 1,247 votes · Scale 1,050–1,800, 0 to 10 · LMArena WebDev arena (step-5-preview-high) · Sep 25, 2026

  • Codebase contextw15—

    No independent measurement yet

Price & availability

Measured the same way, deliberately kept out of the score: what a model costs doesn’t change what it can do.

  • Cost efficiencynot scored—

    No independent measurement yet

w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works

More in Coding & Agentic

View all →
OpenAI
Under evaluation
  • Excellent terminal automation, git operations, and CI/CD debugging
  • Far more token-efficient than reasoning-heavy rivals on routine tasks
  • Needs detailed, unambiguous instructions; struggles with vague requests
  • Smaller context window than some rivals, a constraint on huge monorepos
Anthropic
  • Strong at inferring intent from vague prompts and architectural context
  • 1M-token context supports coherent multi-file, cross-repo refactors
  • Can use many more tokens than leaner coding models on routine work
  • Narrates its reasoning at length, which slows down quick tasks
Alibaba
  • Open-weight performance within striking distance of proprietary leaders
  • Free to self-host, appealing for cost-sensitive or air-gapped teams
  • Requires serious infrastructure to run at full size
  • Tooling and IDE integrations are less mature than Copilot, Cursor, or Codex
Cursor
Under evaluation
  • Tight, real-time feedback loop; you see and steer every change
  • Affordable flat-rate pricing for all-day assistance
  • Needs a developer actively driving; not built for unattended runs
  • Background and async agent mode is still early and limited