Models / Coding & Agentic Claude Haiku 4.5 Claude Haiku 4.5 5.5 /10
Same model as Claude Haiku 4.5 in Language Models, listed again here because it’s used as a coding model and scored on that rubric. It counts once in the directory’s totals.
Maker Anthropic
Released Oct 1, 2025 Strengths & weaknesses + Resolves 66.6% of SWE-bench Verified issues, unusually high for a small, cheap tier + Anthropic’s cheapest tier with tool calling and structured outputs, which suits high-volume automated fixes − The lowest SWE-bench figure among the scored coding models here − A 200k-token context caps how much of a large codebase it can hold at once
Evaluation 100% of weight measured Scored on Claude Haiku 4.5 · sources as of Sep 20, 2026
Confidence band 5.5–5.5 · a model inside this range isn’t meaningfully apart from this one
Code qualityw45 6.7
Rule A 66.6% resolved · 66.6% ÷ 10 · SWE-bench Verified, % resolved (mini-SWE-agent harness) (Claude 4.5 Haiku) · Feb 17, 2026
Web developmentw40 3.7
Rule B 1,329 rating (95% 1,324–1,334) from 27,785 votes · Scale 1,050–1,800, 0 to 10 · LMArena WebDev arena (claude-haiku-4-5-20251001) · Sep 11, 2026
Codebase contextw15 6.6
Rule B 200,000 tokens · Scale 8,192–1,048,576 tokens, 0 to 10 · OpenRouter, context window · Sep 20, 2026
Price & availability Measured the same way, deliberately kept out of the score: what a model costs doesn’t change what it can do.
w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works
Enterprise fit
Highest-ROI use cases Where Claude Haiku 4.5 fits inside a company, ranked by the strength of published ROI evidence for each use case.
All 100 enterprise use cases → 01 IT & engineering ▸ Built for this kind of work (Coding & Agentic)02 Cross-functional operations ▸ Built for this kind of work (Coding & Agentic)
+ Excellent terminal automation, git operations, and CI/CD debugging + Far more token-efficient than reasoning-heavy rivals on routine tasks − Needs detailed, unambiguous instructions; struggles with vague requests − Smaller context window than some rivals, a constraint on huge monorepos
+ Strong at inferring intent from vague prompts and architectural context + 1M-token context supports coherent multi-file, cross-repo refactors − Can use many more tokens than leaner coding models on routine work − Narrates its reasoning at length, which slows down quick tasks
+ Open-weight performance within striking distance of proprietary leaders + Free to self-host, appealing for cost-sensitive or air-gapped teams − Requires serious infrastructure to run at full size − Tooling and IDE integrations are less mature than Copilot, Cursor, or Codex
+ Tight, real-time feedback loop; you see and steer every change + Affordable flat-rate pricing for all-day assistance − Needs a developer actively driving; not built for unattended runs − Background and async agent mode is still early and limited