Section 1 of 7 — 30 models

Language Models

General-purpose chat and reasoning models — the models most people mean when they say “AI.” Frontier flagships, their budget tiers, and the open-weight and regional alternatives.

OpenAI
7.2/10
  • Best-in-class agentic reasoning, with search, code execution, and computer use in one API
  • Disciplined, low-hallucination output on long, multi-step tasks
  • Long-context requests above roughly 272K tokens get repriced sharply higher
  • Slower to produce a first answer than most rivals at max reasoning
Anthropic
  • Anthropic's recommended model for complex, high-stakes work, with strong reasoning-to-cost
  • Zero-data-retention eligible, useful for regulated or enterprise deployments
  • Sits below the flagship tier in branding despite strong practical scores
  • Costs meaningfully more per token than the Sonnet tier for everyday tasks
Anthropic
  • Fast and capable, priced for everyday production use
  • Strong default for coding and long documents without Opus-level cost
  • Enabling maximum thinking mode can quietly balloon token spend
  • Trails Opus on the hardest multi-step reasoning problems
Anthropic
  • Cheapest Claude tier, well suited to high-volume, simple tasks
  • Low latency, good for chat-style and classification workloads
  • Smaller context window (200K tokens) than Sonnet 5 and Opus 5 (1M)
  • Noticeably weaker on hard reasoning than Sonnet or Opus
Google
  • Native multimodal input across text, image, audio, video, and PDFs
  • Very large context window with strong long-document recall
  • Slower first response than Google's own Flash tier
  • No native image or audio output, understanding only
Google
  • Among the fastest decoding speeds of any frontier-class model
  • Aggressively priced for high-throughput production traffic
  • Newer Flash versions (3.5 to 3.8) have since replaced it as Google’s default fast model
  • Higher reasoning settings add noticeable cost and latency
xAI
6.1/10
  • Leads several agentic and terminal-automation benchmarks
  • Native, real-time X and web search built into the model
  • Reasoning mode cannot be turned off, adding cost and latency to every call
  • Hits a steep pricing cliff once context passes roughly 200K tokens
DeepSeek
  • Genuinely open-weight frontier model you can self-host, unlike closed rivals
  • Very low API cost, especially during off-peak pricing windows
  • V4 Pro is text-only; image input is limited to an experimental Flash variant
  • Trails the closed frontier labs on aggregate intelligence benchmarks
Zhipu AI
7.6/10
  • Coding performance that rivals or matches Claude on several benchmarks
  • Open-weight, giving flexibility for self-hosted or fine-tuned deployments
  • Smaller ecosystem of tooling and integrations than the big three labs
  • Less battle-tested in production outside China-based deployments
Alibaba
5.1/10
  • Competitive with top closed models on reasoning and coding benchmarks
  • Broad family of sizes, from edge-friendly to frontier-scale
  • Closed weights: unlike Qwen’s open models, it can’t be self-hosted
  • English-language polish lags slightly behind Western frontier labs
Meta
6.3/10
  • The most widely deployed open-weight model in enterprise settings
  • Large ecosystem of fine-tunes, tooling, and community support
  • License restricts free use once a deployer passes 700M monthly users
  • Falls behind Qwen, DeepSeek, and GLM on several 2026 benchmarks
Mistral AI
  • Strong multilingual performance, especially across European languages
  • Apache 2.0 license permits unrestricted commercial use
  • Smaller research budget than the US/China frontier labs shows in ceiling performance
  • Less name recognition can complicate enterprise procurement
Google
7.7/10
  • Purpose-built for on-device and edge deployment
  • Small variants, from 4B to 12B, run well on modest hardware
  • Not intended to compete with frontier models on hard reasoning
  • Limited context window compared to Gemini's cloud models
Microsoft
4.9/10
  • Punches above its weight for its size, especially on math and logic
  • Small enough to run locally or at the edge cheaply
  • Narrower general-knowledge breadth than larger frontier models
  • Less suited to long, open-ended agentic tasks
Mistral AI
  • One open-weight model that combines reasoning, image input, and agentic coding, under Apache 2.0
  • Mixture-of-experts design activates only 6.5B of 119B parameters per token, keeping inference efficient
  • A small-tier model: trails frontier flagships on the hardest reasoning tasks
  • At 119B total parameters, self-hosting needs far more memory than earlier Small models
Moonshot AI
5.1/10
  • Excels at long documents, source-heavy writing, and research briefs, holding context well
  • Open-weight with strong agentic and coding chops in its K2 update
  • Smaller ecosystem and fewer integrations than the big three closed labs
  • Less proven for latency-sensitive, real-time production use
Baidu
  • Deep integration with Baidu's search and China-market ecosystem
  • Strong Mandarin-language understanding and generation
  • Weaker mindshare and tooling support outside China
  • English-language performance trails the Western frontier labs
TII (UAE)
  • Backed by the UAE's Technology Innovation Institute, with strong regional language coverage including Arabic
  • Open weights under a permissive license, good for sovereign or on-premise deployments
  • Trails the largest US/China labs on raw benchmark ceiling
  • Smaller third-party tooling ecosystem than Llama or Qwen
NVIDIA
  • Fully open: weights, training data, and recipes released under the permissive OpenMDW-1.1 licence
  • 550B-parameter hybrid Mamba-Transformer (55B active) built for long-running agents
  • At 550B parameters, self-hosting needs a multi-GPU cluster
  • Its most efficient format (NVFP4) is designed for NVIDIA hardware
AI21 Labs
  • Hybrid Transformer-Mamba architecture gives very high throughput and low memory use at long context
  • One of the largest context windows among open-weight models, aimed at enterprise long-document work
  • Needs its own proprietary quantization approach to hit those efficiency numbers
  • Much smaller track record and mindshare than Llama, Qwen, or DeepSeek
Tencent
  • Strong presence in the Chinese consumer and enterprise market via Tencent's ecosystem
  • Broad multimodal family spanning text, image, and video under one brand
  • Limited adoption and tooling outside China
  • Less benchmarked against Western frontier models in independent tests
Reka AI
  • Compact, efficient multimodal model that's easy to self-host
  • Good cost-to-performance ratio for mid-tier multimodal tasks
  • Trails frontier labs on the hardest reasoning benchmarks
  • Much smaller brand recognition and community than the major labs
xAI
  • Cheaper, faster budget tier of Grok for high-volume or latency-sensitive use
  • Still inherits Grok's native X and web search integration
  • Clearly less capable than Grok 4.6 on hard reasoning
  • Superseded by Grok 4.1 Fast and newer fast models from xAI
ByteDance
  • Deeply integrated into ByteDance's own consumer apps and ecosystem
  • Competitive general-purpose performance at low cost
  • Primarily built for the Chinese market, with limited presence elsewhere
  • Less transparent benchmarking against global frontier models
Upstage
  • Strong performance for its size, popular in Korean-language and enterprise RAG use cases
  • Efficient enough to run cost-effectively at scale
  • Narrower global mindshare than the major US/China labs
  • Falls behind frontier models on the hardest general reasoning tasks
01.AI
  • Fast, low-cost model that performs well for its price point
  • Reasonable multilingual coverage including English and Chinese
  • Trails top-tier frontier models on complex reasoning and coding
  • Smaller developer ecosystem than Qwen or DeepSeek
Cohere
  • Purpose-built for broad multilingual coverage across dozens of languages, including many under-served ones
  • Open-weight, useful for research and localization-heavy applications
  • Not intended to compete with frontier models on English-centric reasoning
  • Smaller production track record than Cohere's commercial Command line
OpenAI
4.0/10
  • Mature, well-tested legacy model still widely integrated across products
  • Good balance of speed, cost, and multimodal input for everyday use
  • Clearly superseded by GPT-5.6 on hard reasoning and agentic tasks
  • OpenAI's roadmap increasingly steers new users toward newer models
Cohere
  • Built specifically for retrieval-augmented generation and enterprise search workflows
  • Strong citation and grounding behavior when paired with a document store
  • Less suited to open-ended creative or general chat use
  • Now sits below Cohere's newer Command A+ flagship
Huawei
  • Backed by Huawei's own chip and cloud stack, tuned for performance on Huawei's Ascend hardware
  • Positioned for China's state and enterprise sector, with sovereign-cloud appeal
  • Very limited availability and independent benchmarking outside China
  • Smaller developer community and third-party tooling than rivals