LOADING
PLEASE WAIT...
LOADING
PLEASE WAIT...
Compare the most popular language models side by side: context window, max output tokens and reference API pricing per 1M tokens. Filter by provider and pin models to compare.
| MODEL | PROVIDER | CONTEXT | MAX OUT | INPUT $/1M | OUTPUT $/1M | |
|---|---|---|---|---|---|---|
| Claude 3.7 Sonnet Coding, long outputs (legacy) | Anthropic | 200K | 66K | $3.00 | $15.00 | |
| Claude Haiku 4 Fast, cheap everyday tasks | Anthropic | 200K | 66K | $1.00 | $5.00 | |
| Claude Opus 4 Hardest tasks, deep analysis | Anthropic | 200K | 33K | $15.00 | $75.00 | |
| Claude Sonnet 4 Coding, agents, balanced quality | Anthropic | 200K | 66K | $3.00 | $15.00 | |
| DeepSeek-R1 Reasoning, open weights | DeepSeek | 131K | 16K | $0.55 | $2.19 | |
| DeepSeek-V3 Cheap general + coding | DeepSeek | 131K | 16K | $0.27 | $1.10 | |
| Gemini 2.5 Flash Fast, cheap, big context | 1.0M | 66K | $0.30 | $2.50 | ||
| Gemini 2.5 Flash-Lite Cheapest high-volume tasks | 1.0M | 66K | $0.10 | $0.40 | ||
| Gemini 2.5 Pro Long context, multimodal reasoning | 1.0M | 66K | $1.25 | $10.00 | ||
| GPT-4.1 Long context, coding, agents | OpenAI | 1.0M | 33K | $2.00 | $8.00 | |
| GPT-4.1 mini Long context at low cost | OpenAI | 1.0M | 33K | $0.40 | $1.60 | |
| GPT-4o General assistant, vision, everyday tasks | OpenAI | 128K | 16K | $2.50 | $10.00 | |
| GPT-4o mini Cheap fast tasks, classification, chat | OpenAI | 128K | 16K | $0.15 | $0.60 | |
| Grok 4 Real-time data, X integration | xAI | 262K | 33K | $3.00 | $15.00 | |
| Llama 3.3 70B Open weights, self-hosted | Meta | 131K | 8K | $0.25 | $0.70 | |
| Llama 4 Maverick Open weights, huge context | Meta | 1.0M | 128K | $0.20 | $0.60 | |
| Mistral Large 2 Multilingual, enterprise | Mistral | 131K | 8K | $2.00 | $6.00 | |
| Mistral Small Fast lightweight tasks | Mistral | 131K | 8K | $0.20 | $0.60 | |
| o3 Hard reasoning, math, code | OpenAI | 200K | 100K | $10.00 | $40.00 | |
| o3-mini Reasoning on a budget | OpenAI | 200K | 100K | $1.10 | $4.40 |
No. Prices are reference values in USD per 1M tokens collected from public provider pages and change often — some providers also offer batch discounts and cached-input pricing. Always confirm on the official pricing page before budgeting. The cheapest visible input and output prices are highlighted in the table.
This comparison is a static table — it runs entirely in your browser with no tracking and no data sent anywhere. No model names or prices are transmitted.
There is no single winner — it depends on the task. For everyday assistant work, mid-size models like GPT-4o, Claude Sonnet and Gemini 2.5 Flash are a great balance of quality, speed and price. For complex reasoning, frontier models like o3, Claude Opus and Gemini 2.5 Pro lead the benchmarks.
It is the maximum amount of text the model can consider in one request, measured in tokens (roughly 4 characters each). 128K ≈ a 300-page book; 1M ≈ multiple books. Bigger is not always better — cost and latency grow with the context you send.
Input tokens are everything you send (system prompt, messages, documents). Output tokens are the model's answer. Output is usually priced higher because generating tokens is more expensive for providers.
No. They are reference values in USD per 1M tokens collected from public pricing pages. Models, prices and limits change frequently, so always check the provider's official page before estimating costs.
Yes — that is the point. Pin up to 4 models (any mix of OpenAI, Anthropic, Google, Meta, DeepSeek, Mistral and xAI) and they appear side by side in the strip above the table.
Yes. It is a static table in your browser — no tracking, no requests, no data sent anywhere.