OpenRouter Models Compared
Which AI model should you use? Compare context windows, pricing per million tokens, modality support and best-fit workloads across the OpenRouter ecosystem.
Updated: 2026-09-05
Best overall
Qwen3.8 Max (0902)
Alibaba
Best for coding
Claude Sonnet 5
Anthropic
Best for long context
Xiaomi MiMo-V2.5
formerly Hunter Alpha
Best for budget
Xiaomi MiMo-V2.5
formerly Hunter Alpha
Best multimodal
Xiaomi MiMo-V2.5
formerly Hunter Alpha
Best for agents
DeepSeek V4 Pro
DeepSeek
All models in this comparison
| Model | Provider | Context | Input / Output / 1M | Modalities | Best for |
|---|---|---|---|---|---|
| Xiaomi MiMo-V2.5 xiaomi/mimo-v2.5 formerly Hunter Alpha | Xiaomi | 1.05M tokens | $0.140 / $0.280 | Text, Vision, Audio, Video | Long Context, Budget, Multimodal |
| Z.ai GLM 5.3 Flash z-ai/glm-5.3-flash formerly OX Alpha | Z.ai | 1.31M tokens | $0.075 / $0.250 | Text, Vision, Video | Budget, Long Context, Multimodal |
| DeepSeek V4 Flash deepseek/deepseek-v4-flash-0731 | DeepSeek | 1.31M tokens | $0.065 / $0.180 | Text | Budget, Long Context |
| DeepSeek V4 Pro deepseek/deepseek-v4-pro-0813 | DeepSeek | 1.05M tokens | $0.579 / $1.74 | Text | Long Context, Budget, Agents |
| Qwen3.8 Flash qwen/qwen3.8-flash | Alibaba | 1M tokens | $0.150 / $0.470 | Text, Vision, Video | Budget, Long Context, Multimodal |
| Qwen3.8 Max (0902) qwen/qwen3.8-max-0902 | Alibaba | 1M tokens | $2.00 / $6.00 | Text, Vision, Video | Overall, Multimodal, Long Context |
| Google Gemini 3.7 Flash google/gemini-3.7-flash | 1.05M tokens | $0.750 / $3.75 | Text, Vision, Audio, Video, Files | Multimodal, Overall, Long Context | |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | Anthropic | 1M tokens | $2.00 / $10.00 | Text, Vision, Files | Overall, Coding, Agents |
| Claude Opus 5 anthropic/claude-opus-5 | Anthropic | 1M tokens | $5.00 / $25.00 | Text, Vision, Files | Overall, Agents, Coding |
| OpenAI GPT-5.6 Luna openai/gpt-5.6-luna | OpenAI | 1.05M tokens | $0.200 / $1.20 | Text, Vision, Files | Budget, Coding, Agents |
| OpenAI GPT-5.6 Sol openai/gpt-5.6-sol | OpenAI | 1.05M tokens | $2.00 / $10.00 | Text, Vision, Files | Overall, Coding, Agents |
| OpenAI GPT-5.6 Terra openai/gpt-5.6-terra | OpenAI | 1.05M tokens | $2.00 / $12.00 | Text, Vision, Files | Overall, Multimodal, Agents |
| Mistral Medium 3.5 mistralai/mistral-medium-3-5 | Mistral AI | 262K tokens | $1.50 / $7.50 | Text, Vision, Files | Overall, Coding |
| Meta Llama 4 Maverick meta-llama/llama-4-maverick | Meta | 1.05M tokens | $0.200 / $0.696 | Text, Vision | Budget, Long Context, Multimodal |
| Cohere Command A cohere/command-a | Cohere | 256K tokens | $2.50 / $10.00 | Text | Agents, Coding |
Pricing is shown as input/output per 1M tokens and is based on a 2026-09-05 snapshot. Provider limits and surcharges can change.
Best picks by workload
Best overall
- 1.
Qwen3.8 Max (0902)
Broad multimodal coverage
- 2.
Google Gemini 3.7 Flash
Very broad multimodal input support
- 3.
Claude Sonnet 5
Strong coding, writing and long-context reasoning
Best for coding
- 1.
Claude Sonnet 5
Strong coding, writing and long-context reasoning
- 2.
Claude Opus 5
Highest reasoning quality in the snapshot
- 3.
OpenAI GPT-5.6 Luna
Fast, cost-efficient reasoning
Best for long context
- 1.
Xiaomi MiMo-V2.5
1.05M-token context window at a very low price
- 2.
Z.ai GLM 5.3 Flash
Very low input and output pricing
- 3.
DeepSeek V4 Flash
Extremely low input and output cost
Best for budget
- 1.
Xiaomi MiMo-V2.5
1.05M-token context window at a very low price
- 2.
Z.ai GLM 5.3 Flash
Very low input and output pricing
- 3.
DeepSeek V4 Flash
Extremely low input and output cost
Best multimodal
- 1.
Xiaomi MiMo-V2.5
1.05M-token context window at a very low price
- 2.
Z.ai GLM 5.3 Flash
Very low input and output pricing
- 3.
Qwen3.8 Flash
Low-latency multimodal processing
Best for agents
- 1.
DeepSeek V4 Pro
Balanced cost/performance for long-form reasoning
- 2.
Claude Sonnet 5
Strong coding, writing and long-context reasoning
- 3.
Claude Opus 5
Highest reasoning quality in the snapshot
Head-to-head comparisons
Claude Sonnet 5 vs OpenAI GPT-5.6 Sol
Claude Sonnet 5 has a lower output price, while GPT-5.6 Sol is positioned as a balanced frontier option with strong tooling.
Claude Opus 5 vs OpenAI GPT-5.6 Terra
Claude Opus 5 has lower output pricing than GPT-5.6 Terra, while Terra emphasizes strong multimodal and file reasoning.
DeepSeek V4 Flash vs Z.ai GLM 5.3 Flash
GLM 5.3 Flash adds vision and video input, while DeepSeek V4 Flash is text-only but slightly cheaper and still very large-context.
Xiaomi MiMo-V2.5 vs DeepSeek V4 Flash
Xiaomi MiMo-V2.5 supports multimodal input; DeepSeek V4 Flash is text-only but has a lower price and comparable large-context reach.
Xiaomi MiMo-V2.5 vs Z.ai GLM 5.3 Flash
GLM 5.3 Flash has a larger context window and lower pricing; MiMo-V2.5 adds audio input and may perform differently on long-context quality.
Qwen3.8 Flash vs Google Gemini 3.7 Flash
Qwen3.8 Flash is significantly cheaper and supports image/video, while Gemini 3.7 Flash adds audio and files with a broader multimodal surface.
Google Gemini 3.7 Flash vs Claude Sonnet 5
Gemini 3.7 Flash offers broader multimodal input at lower cost; Claude Sonnet 5 is usually stronger for code, writing and agent reliability.
DeepSeek V4 Pro vs Claude Opus 5
Claude Opus 5 is the premium reasoning option, while DeepSeek V4 Pro offers a 1M-token text workflow at a much lower price.
Meta Llama 4 Maverick vs DeepSeek V4 Flash
Both are low-cost large-context options; Llama 4 Maverick adds vision, while DeepSeek V4 Flash is text-only with lower output pricing.
OpenAI GPT-5.6 Luna vs Qwen3.8 Flash
GPT-5.6 Luna has a slight edge on context and tooling; Qwen3.8 Flash is cheaper and supports video input.
Frequently asked questions
What is the best model on OpenRouter for coding?
Claude Sonnet 5, Claude Opus 5 and GPT-5.6 Sol are strong general-purpose choices for coding. Sonnet 5 is often the best price/quality default; Opus 5 is better for difficult architecture or debugging work.
Which OpenRouter model has the largest context window?
In the current snapshot, DeepSeek V4 Flash and GLM 5.3 Flash offer about 1.31M tokens, with several other models around 1M tokens. Always verify provider-specific limits before long-context production use.
What is the cheapest model on OpenRouter?
DeepSeek V4 Flash, GLM 5.3 Flash and Llama 4 Maverick are among the lowest-cost options. For free routes, see our free models page; some free models have rate limits and smaller context windows.
Are there free models on OpenRouter?
Yes. OpenRouter exposes some models with a :free suffix. These are useful for testing and low-volume workloads, but rate limits, context limits and quality vary significantly.
What happened to Hunter Alpha and OX Alpha?
Hunter Alpha is now known as Xiaomi MiMo-V2.5 and OX Alpha was revealed as GLM 5.3 Flash. We keep those names as historical aliases, but no longer treat them as primary SEO targets.
How do OpenRouter models compare to ChatGPT?
OpenRouter is not a single model; it is a model router and marketplace. The best comparison depends on the model you choose, the input modality, context length and pricing tier.