Reka AI

Reka Flash 3

Maker
Reka AI
Origin
United States
Released
Mar 2025

Strengths & weaknesses

  • Compact, efficient multimodal model that's easy to self-host
  • Good cost-to-performance ratio for mid-tier multimodal tasks
  • Trails frontier labs on the hardest reasoning benchmarks
  • Much smaller brand recognition and community than the major labs

Evaluation

In assessment · 40% measured

Scored on Reka Flash 3 · sources as of Sep 19, 2026

  • Reasoningw30

    No independent measurement yet

  • Accuracyw20

    No independent measurement yet

  • Writingw10

    No independent measurement yet

  • Long contextw101.2

    Rule B65,536 tokens · Ranks 16 of 17 measured in Language Models · OpenRouter, context window · Sep 19, 2026

  • Cost efficiencyw209.4

    Rule B$0.13 per 1M tokens (blended) · Ranks 2 of 17 measured in Language Models · OpenRouter, blended price per token · Sep 19, 2026

  • Deployabilityw104.0

    Rule C2 of 5 items · 2 × 2 points · Deployability checklist (OpenRouter data) · Sep 19, 2026

w = weight, each criterion’s share of the overall score. Missing marks don’t count for or against. How scoring works

Enterprise fit

Highest-ROI use cases

Where Reka Flash 3 fits inside a company, ranked by the strength of published ROI evidence for each use case.

All 100 enterprise use cases →
  1. 01
    Strategy, consulting, finance
    • Built for this kind of work (Language Models)
  2. 02
    Marketing & sales
    • Built for this kind of work (Language Models)
  3. 03
    Customer supportField study
    Customer service
    • Built for this kind of work (Language Models)

More in Language Models

View all →
OpenAI
7.2/10
  • Best-in-class agentic reasoning, with search, code execution, and computer use in one API
  • Disciplined, low-hallucination output on long, multi-step tasks
  • Long-context requests above roughly 272K tokens get repriced sharply higher
  • Slower to produce a first answer than most rivals at max reasoning
Anthropic
  • Anthropic's recommended model for complex, high-stakes work, with strong reasoning-to-cost
  • Zero-data-retention eligible, useful for regulated or enterprise deployments
  • Sits below the flagship tier in branding despite strong practical scores
  • Costs meaningfully more per token than the Sonnet tier for everyday tasks
Anthropic
  • Fast and capable, priced for everyday production use
  • Strong default for coding and long documents without Opus-level cost
  • Enabling maximum thinking mode can quietly balloon token spend
  • Trails Opus on the hardest multi-step reasoning problems
Anthropic
  • Cheapest Claude tier, well suited to high-volume, simple tasks
  • Low latency, good for chat-style and classification workloads
  • Smaller context window (200K tokens) than Sonnet 5 and Opus 5 (1M)
  • Noticeably weaker on hard reasoning than Sonnet or Opus