OpenRouter Models Compared

Which AI model should you use? Compare context windows, pricing per million tokens, modality support and best-fit workloads across the OpenRouter ecosystem.

Updated: 2026-09-05

Best overall

Qwen3.8 Max (0902)

Alibaba

Best for coding

Claude Sonnet 5

Anthropic

Best for long context

Xiaomi MiMo-V2.5

formerly Hunter Alpha

Best for budget

Xiaomi MiMo-V2.5

formerly Hunter Alpha

Best multimodal

Xiaomi MiMo-V2.5

formerly Hunter Alpha

Best for agents

DeepSeek V4 Pro

DeepSeek

All models in this comparison

ModelProviderContextInput / Output / 1MModalitiesBest for
Xiaomi MiMo-V2.5
xiaomi/mimo-v2.5
formerly Hunter Alpha
Xiaomi1.05M tokens$0.140 / $0.280Text, Vision, Audio, VideoLong Context, Budget, Multimodal
Z.ai GLM 5.3 Flash
z-ai/glm-5.3-flash
formerly OX Alpha
Z.ai1.31M tokens$0.075 / $0.250Text, Vision, VideoBudget, Long Context, Multimodal
DeepSeek V4 Flash
deepseek/deepseek-v4-flash-0731
DeepSeek1.31M tokens$0.065 / $0.180TextBudget, Long Context
DeepSeek V4 Pro
deepseek/deepseek-v4-pro-0813
DeepSeek1.05M tokens$0.579 / $1.74TextLong Context, Budget, Agents
Qwen3.8 Flash
qwen/qwen3.8-flash
Alibaba1M tokens$0.150 / $0.470Text, Vision, VideoBudget, Long Context, Multimodal
Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902
Alibaba1M tokens$2.00 / $6.00Text, Vision, VideoOverall, Multimodal, Long Context
Google Gemini 3.7 Flash
google/gemini-3.7-flash
Google1.05M tokens$0.750 / $3.75Text, Vision, Audio, Video, FilesMultimodal, Overall, Long Context
Claude Sonnet 5
anthropic/claude-sonnet-5
Anthropic1M tokens$2.00 / $10.00Text, Vision, FilesOverall, Coding, Agents
Claude Opus 5
anthropic/claude-opus-5
Anthropic1M tokens$5.00 / $25.00Text, Vision, FilesOverall, Agents, Coding
OpenAI GPT-5.6 Luna
openai/gpt-5.6-luna
OpenAI1.05M tokens$0.200 / $1.20Text, Vision, FilesBudget, Coding, Agents
OpenAI GPT-5.6 Sol
openai/gpt-5.6-sol
OpenAI1.05M tokens$2.00 / $10.00Text, Vision, FilesOverall, Coding, Agents
OpenAI GPT-5.6 Terra
openai/gpt-5.6-terra
OpenAI1.05M tokens$2.00 / $12.00Text, Vision, FilesOverall, Multimodal, Agents
Mistral Medium 3.5
mistralai/mistral-medium-3-5
Mistral AI262K tokens$1.50 / $7.50Text, Vision, FilesOverall, Coding
Meta Llama 4 Maverick
meta-llama/llama-4-maverick
Meta1.05M tokens$0.200 / $0.696Text, VisionBudget, Long Context, Multimodal
Cohere Command A
cohere/command-a
Cohere256K tokens$2.50 / $10.00TextAgents, Coding

Pricing is shown as input/output per 1M tokens and is based on a 2026-09-05 snapshot. Provider limits and surcharges can change.

Best picks by workload

Best overall

  • 1.

    Qwen3.8 Max (0902)

    Broad multimodal coverage

  • 2.

    Google Gemini 3.7 Flash

    Very broad multimodal input support

  • 3.

    Claude Sonnet 5

    Strong coding, writing and long-context reasoning

Best for coding

  • 1.

    Claude Sonnet 5

    Strong coding, writing and long-context reasoning

  • 2.

    Claude Opus 5

    Highest reasoning quality in the snapshot

  • 3.

    OpenAI GPT-5.6 Luna

    Fast, cost-efficient reasoning

Best for long context

  • 1.

    Xiaomi MiMo-V2.5

    1.05M-token context window at a very low price

  • 2.

    Z.ai GLM 5.3 Flash

    Very low input and output pricing

  • 3.

    DeepSeek V4 Flash

    Extremely low input and output cost

Best for budget

  • 1.

    Xiaomi MiMo-V2.5

    1.05M-token context window at a very low price

  • 2.

    Z.ai GLM 5.3 Flash

    Very low input and output pricing

  • 3.

    DeepSeek V4 Flash

    Extremely low input and output cost

Best multimodal

  • 1.

    Xiaomi MiMo-V2.5

    1.05M-token context window at a very low price

  • 2.

    Z.ai GLM 5.3 Flash

    Very low input and output pricing

  • 3.

    Qwen3.8 Flash

    Low-latency multimodal processing

Best for agents

  • 1.

    DeepSeek V4 Pro

    Balanced cost/performance for long-form reasoning

  • 2.

    Claude Sonnet 5

    Strong coding, writing and long-context reasoning

  • 3.

    Claude Opus 5

    Highest reasoning quality in the snapshot

Head-to-head comparisons

Claude Sonnet 5 vs OpenAI GPT-5.6 Sol

Claude Sonnet 5 has a lower output price, while GPT-5.6 Sol is positioned as a balanced frontier option with strong tooling.

Claude Opus 5 vs OpenAI GPT-5.6 Terra

Claude Opus 5 has lower output pricing than GPT-5.6 Terra, while Terra emphasizes strong multimodal and file reasoning.

DeepSeek V4 Flash vs Z.ai GLM 5.3 Flash

GLM 5.3 Flash adds vision and video input, while DeepSeek V4 Flash is text-only but slightly cheaper and still very large-context.

Xiaomi MiMo-V2.5 vs DeepSeek V4 Flash

Xiaomi MiMo-V2.5 supports multimodal input; DeepSeek V4 Flash is text-only but has a lower price and comparable large-context reach.

Xiaomi MiMo-V2.5 vs Z.ai GLM 5.3 Flash

GLM 5.3 Flash has a larger context window and lower pricing; MiMo-V2.5 adds audio input and may perform differently on long-context quality.

Qwen3.8 Flash vs Google Gemini 3.7 Flash

Qwen3.8 Flash is significantly cheaper and supports image/video, while Gemini 3.7 Flash adds audio and files with a broader multimodal surface.

Google Gemini 3.7 Flash vs Claude Sonnet 5

Gemini 3.7 Flash offers broader multimodal input at lower cost; Claude Sonnet 5 is usually stronger for code, writing and agent reliability.

DeepSeek V4 Pro vs Claude Opus 5

Claude Opus 5 is the premium reasoning option, while DeepSeek V4 Pro offers a 1M-token text workflow at a much lower price.

Meta Llama 4 Maverick vs DeepSeek V4 Flash

Both are low-cost large-context options; Llama 4 Maverick adds vision, while DeepSeek V4 Flash is text-only with lower output pricing.

OpenAI GPT-5.6 Luna vs Qwen3.8 Flash

GPT-5.6 Luna has a slight edge on context and tooling; Qwen3.8 Flash is cheaper and supports video input.

Frequently asked questions

What is the best model on OpenRouter for coding?

Claude Sonnet 5, Claude Opus 5 and GPT-5.6 Sol are strong general-purpose choices for coding. Sonnet 5 is often the best price/quality default; Opus 5 is better for difficult architecture or debugging work.

Which OpenRouter model has the largest context window?

In the current snapshot, DeepSeek V4 Flash and GLM 5.3 Flash offer about 1.31M tokens, with several other models around 1M tokens. Always verify provider-specific limits before long-context production use.

What is the cheapest model on OpenRouter?

DeepSeek V4 Flash, GLM 5.3 Flash and Llama 4 Maverick are among the lowest-cost options. For free routes, see our free models page; some free models have rate limits and smaller context windows.

Are there free models on OpenRouter?

Yes. OpenRouter exposes some models with a :free suffix. These are useful for testing and low-volume workloads, but rate limits, context limits and quality vary significantly.

What happened to Hunter Alpha and OX Alpha?

Hunter Alpha is now known as Xiaomi MiMo-V2.5 and OX Alpha was revealed as GLM 5.3 Flash. We keep those names as historical aliases, but no longer treat them as primary SEO targets.

How do OpenRouter models compare to ChatGPT?

OpenRouter is not a single model; it is a model router and marketplace. The best comparison depends on the model you choose, the input modality, context length and pricing tier.