Section 1 of 7 — 30 models
Language Models Language Models General-purpose chat and reasoning models — the models most people mean when they say “AI.” Frontier flagships, their budget tiers, and the open-weight and regional alternatives.
+ Best-in-class agentic reasoning, with search, code execution, and computer use in one API + Disciplined, low-hallucination output on long, multi-step tasks − Long-context requests above roughly 272K tokens get repriced sharply higher − Slower to produce a first answer than most rivals at max reasoning
+ Anthropic's recommended model for complex, high-stakes work, with strong reasoning-to-cost + Zero-data-retention eligible, useful for regulated or enterprise deployments − Sits below the flagship tier in branding despite strong practical scores − Costs meaningfully more per token than the Sonnet tier for everyday tasks
+ Fast and capable, priced for everyday production use + Strong default for coding and long documents without Opus-level cost − Enabling maximum thinking mode can quietly balloon token spend − Trails Opus on the hardest multi-step reasoning problems
+ Cheapest Claude tier, well suited to high-volume, simple tasks + Low latency, good for chat-style and classification workloads − Smaller context window (200K tokens) than Sonnet 5 and Opus 5 (1M) − Noticeably weaker on hard reasoning than Sonnet or Opus
+ Native multimodal input across text, image, audio, video, and PDFs + Very large context window with strong long-document recall − Slower first response than Google's own Flash tier − No native image or audio output, understanding only
+ Among the fastest decoding speeds of any frontier-class model + Aggressively priced for high-throughput production traffic − Newer Flash versions (3.5 to 3.8) have since replaced it as Google’s default fast model − Higher reasoning settings add noticeable cost and latency
+ Leads several agentic and terminal-automation benchmarks + Native, real-time X and web search built into the model − Reasoning mode cannot be turned off, adding cost and latency to every call − Hits a steep pricing cliff once context passes roughly 200K tokens
+ Genuinely open-weight frontier model you can self-host, unlike closed rivals + Very low API cost, especially during off-peak pricing windows − V4 Pro is text-only; image input is limited to an experimental Flash variant − Trails the closed frontier labs on aggregate intelligence benchmarks
+ Coding performance that rivals or matches Claude on several benchmarks + Open-weight, giving flexibility for self-hosted or fine-tuned deployments − Smaller ecosystem of tooling and integrations than the big three labs − Less battle-tested in production outside China-based deployments
+ Competitive with top closed models on reasoning and coding benchmarks + Broad family of sizes, from edge-friendly to frontier-scale − Closed weights: unlike Qwen’s open models, it can’t be self-hosted − English-language polish lags slightly behind Western frontier labs
+ The most widely deployed open-weight model in enterprise settings + Large ecosystem of fine-tunes, tooling, and community support − License restricts free use once a deployer passes 700M monthly users − Falls behind Qwen, DeepSeek, and GLM on several 2026 benchmarks
+ Strong multilingual performance, especially across European languages + Apache 2.0 license permits unrestricted commercial use − Smaller research budget than the US/China frontier labs shows in ceiling performance − Less name recognition can complicate enterprise procurement
+ Purpose-built for on-device and edge deployment + Small variants, from 4B to 12B, run well on modest hardware − Not intended to compete with frontier models on hard reasoning − Limited context window compared to Gemini's cloud models
+ Punches above its weight for its size, especially on math and logic + Small enough to run locally or at the edge cheaply − Narrower general-knowledge breadth than larger frontier models − Less suited to long, open-ended agentic tasks
+ One open-weight model that combines reasoning, image input, and agentic coding, under Apache 2.0 + Mixture-of-experts design activates only 6.5B of 119B parameters per token, keeping inference efficient − A small-tier model: trails frontier flagships on the hardest reasoning tasks − At 119B total parameters, self-hosting needs far more memory than earlier Small models
+ Excels at long documents, source-heavy writing, and research briefs, holding context well + Open-weight with strong agentic and coding chops in its K2 update − Smaller ecosystem and fewer integrations than the big three closed labs − Less proven for latency-sensitive, real-time production use
+ Deep integration with Baidu's search and China-market ecosystem + Strong Mandarin-language understanding and generation − Weaker mindshare and tooling support outside China − English-language performance trails the Western frontier labs
+ Backed by the UAE's Technology Innovation Institute, with strong regional language coverage including Arabic + Open weights under a permissive license, good for sovereign or on-premise deployments − Trails the largest US/China labs on raw benchmark ceiling − Smaller third-party tooling ecosystem than Llama or Qwen
+ Fully open: weights, training data, and recipes released under the permissive OpenMDW-1.1 licence + 550B-parameter hybrid Mamba-Transformer (55B active) built for long-running agents − At 550B parameters, self-hosting needs a multi-GPU cluster − Its most efficient format (NVFP4) is designed for NVIDIA hardware
+ Hybrid Transformer-Mamba architecture gives very high throughput and low memory use at long context + One of the largest context windows among open-weight models, aimed at enterprise long-document work − Needs its own proprietary quantization approach to hit those efficiency numbers − Much smaller track record and mindshare than Llama, Qwen, or DeepSeek
+ Strong presence in the Chinese consumer and enterprise market via Tencent's ecosystem + Broad multimodal family spanning text, image, and video under one brand − Limited adoption and tooling outside China − Less benchmarked against Western frontier models in independent tests
+ Compact, efficient multimodal model that's easy to self-host + Good cost-to-performance ratio for mid-tier multimodal tasks − Trails frontier labs on the hardest reasoning benchmarks − Much smaller brand recognition and community than the major labs
+ Cheaper, faster budget tier of Grok for high-volume or latency-sensitive use + Still inherits Grok's native X and web search integration − Clearly less capable than Grok 4.6 on hard reasoning − Superseded by Grok 4.1 Fast and newer fast models from xAI
+ Deeply integrated into ByteDance's own consumer apps and ecosystem + Competitive general-purpose performance at low cost − Primarily built for the Chinese market, with limited presence elsewhere − Less transparent benchmarking against global frontier models
+ Strong performance for its size, popular in Korean-language and enterprise RAG use cases + Efficient enough to run cost-effectively at scale − Narrower global mindshare than the major US/China labs − Falls behind frontier models on the hardest general reasoning tasks
+ Fast, low-cost model that performs well for its price point + Reasonable multilingual coverage including English and Chinese − Trails top-tier frontier models on complex reasoning and coding − Smaller developer ecosystem than Qwen or DeepSeek
+ Purpose-built for broad multilingual coverage across dozens of languages, including many under-served ones + Open-weight, useful for research and localization-heavy applications − Not intended to compete with frontier models on English-centric reasoning − Smaller production track record than Cohere's commercial Command line
+ Mature, well-tested legacy model still widely integrated across products + Good balance of speed, cost, and multimodal input for everyday use − Clearly superseded by GPT-5.6 on hard reasoning and agentic tasks − OpenAI's roadmap increasingly steers new users toward newer models
+ Built specifically for retrieval-augmented generation and enterprise search workflows + Strong citation and grounding behavior when paired with a document store − Less suited to open-ended creative or general chat use − Now sits below Cohere's newer Command A+ flagship
+ Backed by Huawei's own chip and cloud stack, tuned for performance on Huawei's Ascend hardware + Positioned for China's state and enterprise sector, with sovereign-cloud appeal − Very limited availability and independent benchmarking outside China − Smaller developer community and third-party tooling than rivals Next section → Coding & Agentic