Compare LLMs side by side

Which LLM should you use? Put two or three of the 45 tracked models head to head on verified pricing, context window, speed, capabilities and compliance — the best value in every row is highlighted. Click any model name for its full spec sheet. Deterministic, no sign-up, no inference calls — verified against provider docs on 2026-08-16.

Compare LLMs side by side

Compare two models
Anthropic
$2 / $10 per 1M · 1M context
OpenAI
$2 / $12 per 1M · 1.1M context

Claude Sonnet 5 and GPT-5.6 Terra compared across pricing, context, speed, capabilities, and compliance. Claude Sonnet 5 is the cheapest at $2 / $10 per 1M. GPT-5.6 Terra has the largest context window at 1.1M tokens. Claude Sonnet 5 leads on long-context retrieval.

Claude Sonnet 5 vs GPT-5.6 Terra — pricing, context window, speed, capabilities, and compliance side by side. Verified Claude Sonnet 5 2026-08-16, GPT-5.6 Terra 2026-07-31.
Metric
Input price (per 1M)$2$2
Output price (per 1M)$10 (best)$12
Cached input$0.2$0.2
Context window1M tokens1.1M tokens (best)
Max output128k tokens128k tokens
Speed tierbalancedbalanced
Throughput~90 tok/s~105 tok/s (best)
Time to first token~550 ms~550 ms
Prompt cachingYesYes
Extended thinkingYesYes
Vision inputYesYes
Tool useYesYes
Fine-tuningNoYes (best)
Open weightsNoNo
EU regionYesYes
HIPAA eligibleYesYes
Free tierNoNo
Verified2026-08-162026-07-31

Choose Claude Sonnet 5 if…

  • Long-context retrieval is central to your workload

Choose GPT-5.6 Terra if…

  • You want to fine-tune on your own data
Matchups people ask about

Seven questions on use case, volume, latency, and compliance — returns a ranked shortlist with monthly cost estimates. Deterministic, no sign-up, no inference calls.

Best LLM by use case

The picker's answer for a typical buyer in each use case, computed with that use case's benchmark weights and a representative volume. Hit “tune for my volume” on any card to open the quiz and re-rank against your own constraints, or click a model name for its spec sheet. Ranked from 24 current models; the 21 superseded models marked Legacy are still selectable in the comparison above but never recommended here.

Best LLM for coding assistant: Claude Sonnet 5

Code generation, code review, autocomplete, SWE-bench-style work. Current benchmark leader on real-world coding + agents. Claude Sonnet 5 scores 100/100 at $2 / $10 per 1M with 1M context — about $1.2k/mo at 100k calls. Runner-up: Claude Opus 5 (93). Then GLM-5.2 (93).

Best LLM for chatbot: GPT-5.6 Luna

Multi-turn conversations, customer-facing assistants, ticket response. Production-grade chat at $0.20/$1.20 per 1M after the 2026-07-30 price cut. GPT-5.6 Luna scores 93/100 at $0.20 / $1.2 per 1M with 1.1M context — about $611.00/mo at 1M calls. Runner-up: Gemini 3.7 Flash (90). Then Claude Haiku 4.5 (89).

Best LLM for RAG: Gemini 3.1 Pro

Retrieval-augmented generation over your own docs, knowledge bases, wikis. 2M context + top long-context retrieval benchmarks. Gemini 3.1 Pro scores 100/100 at $2 / $12 per 1M with 2M context — about $1.6k/mo at 100k calls. Runner-up: Claude Sonnet 5 (99). Then GPT-5.6 Terra (90).

Best LLM for content writing: Claude Opus 5

Blog posts, emails, ads, scripts, long-form copywriting. Best prose quality and voice adherence. Claude Opus 5 scores 99/100 at $5 / $25 per 1M with 1M context — about $430.00/mo at 10k calls. Runner-up: Claude Sonnet 5 (96). Then GPT-5.6 Terra (93).

Best LLM for data extraction: GPT-5.6 Luna

Extract structured fields, classify text, parse documents to JSON. Strong structured-output reliability at low cost. GPT-5.6 Luna scores 99/100 at $0.20 / $1.2 per 1M with 1.1M context — about $753.00/mo at 1M calls. Runner-up: Claude Haiku 4.5 (91). Then Gemini 3.1 Flash-Lite (88).

Best LLM for autonomous agents: Claude Sonnet 5

Multi-step agents, tool-calling loops, browser agents, coding agents. Best-in-class tool use + long-horizon planning (τ-bench leader). Claude Sonnet 5 scores 100/100 at $2 / $10 per 1M with 1M context — about $2.0k/mo at 100k calls. Runner-up: Claude Opus 5 (94). Then GPT-5.6 Terra (90).

Best LLM for vision: Gemini 3.1 Pro

Image Q&A, document OCR, UI understanding, screenshots. Highest vision benchmark; also handles video natively. Gemini 3.1 Pro scores 100/100 at $2 / $12 per 1M with 2M context — about $880.00/mo at 100k calls. Runner-up: Claude Sonnet 5 (94). Then GPT-5.6 Terra (91).

Best LLM for voice: GPT Realtime 2.1

Realtime voice conversations, call-center agents, IVR replacements. The only mature end-to-end speech-to-speech model in production, now with ~25% lower p95 latency. GPT Realtime 2.1 scores 100/100 at $4 / $24 per 1M with 128k context — about $1.6k/mo at 100k calls. Runner-up: GPT Realtime 2.1 mini (78).

Best LLM for classification: GPT-5.6 Luna

Categorize text, intent detection, toxicity, moderation, sentiment. Reliable for JSON-labeled tasks at scale. GPT-5.6 Luna scores 96/100 at $0.20 / $1.2 per 1M with 1.1M context — about $1.3k/mo at 10M calls. Runner-up: Gemini 3.1 Flash-Lite (95). Then Claude Haiku 4.5 (88).

Best LLM for summarization: Gemini 3.7 Flash

Meeting notes, article TL;DRs, executive summaries, doc digests. 1M context for books and transcripts, at Flash speed. Gemini 3.7 Flash scores 98/100 at $1.5 / $7.5 per 1M with 1.0M context — about $1.1k/mo at 100k calls. Runner-up: Claude Sonnet 5 (90). Then Claude Haiku 4.5 (90).

Best LLM for translation: Qwen3.8-Max

High-fidelity translation, localization, multilingual content. Best for Chinese and pan-Asian languages, 1M context, and open weights since 2026-08-12. Qwen3.8-Max scores 100/100 at $2 / $6 per 1M with 1M context — about $960.00/mo at 100k calls. Runner-up: Gemini 3.1 Pro (95). Then Claude Opus 5 (87).

Best LLM for maps: Gemini 3.1 Pro

Place search, routing, geo reasoning, location-aware assistants, travel. Google knowledge graph + Maps-adjacent training data makes this the default. Gemini 3.1 Pro scores 100/100 at $2 / $12 per 1M with 2M context — about $873.00/mo at 100k calls. Runner-up: Gemini 3.7 Flash (100). Then Claude Sonnet 5 (89).

Best LLM for long documents (100k+ tokens): Claude Sonnet 5

Analyze contracts, books, codebases, long transcripts. 1M context with excellent retrieval + prompt caching. Claude Sonnet 5 scores 100/100 at $2 / $10 per 1M with 1M context — about $742.00/mo at 10k calls. Runner-up: Gemini 3.1 Pro (100). Then Claude Opus 5 (91).

Best LLM for cheap production: Qwen3.7-Flash

Billions of tokens/month where cost dominates. Quality "good enough". Cheapest model in the catalog by ~7x at $0.03/$0.13 per 1M, and still 1M context with vision. Qwen3.7-Flash scores 100/100 at $0.03 / $0.13 per 1M with 1M context — about $630.00/mo at 10M calls. Runner-up: Gemini 3.1 Flash-Lite (96). Then GPT-5.6 Luna (92).

Best LLM for on-prem: GLM-5.2

Self-hosted, private cloud, HIPAA/EU/air-gapped environments. Strongest open-weights coder with usable 1M context; top SWE-Bench Pro. GLM-5.2 scores 96/100 at $1.4 / $4.4 per 1M with 1M context — about $4.3k/mo at 1M calls. Runner-up: Kimi K3 (93). Then DeepSeek V4 Pro (91).

Frequently asked questions about picking an LLM

Quick answers about choosing and comparing LLMs in 2026 — Claude, GPT, Gemini, Grok, Mistral, DeepSeek, GLM, Qwen and Kimi across coding, RAG, agents, vision, maps, and cheap production. Expand any question to read the full answer. Last reviewed 2026-08-16.

How do I choose the right LLM for my use case?

Choose an LLM by filtering on your hardest constraint first — context window, compliance, or latency — then ranking the survivors by benchmarks that match your task and cost at your monthly volume. Relevant benchmarks differ by task: SWE-bench for coding, needle-in-haystack for RAG, MMMU for vision, τ-bench for agents. The LAXIMA LLM Picker automates this in 7 questions and returns a ranked shortlist with monthly cost estimates and the score breakdown behind each pick.

Can I compare two specific LLMs side by side?

Yes — the compare panel at the top of this page puts any two models from the catalog head to head: input and output price, cached input, context window and max output, speed tier, throughput, time to first token, caching, vision, tool use, fine-tuning, open weights, EU and HIPAA flags, free tier, and a data-driven "choose this one if…" verdict for each. The best value in every row is highlighted. The comparison is computed in your browser from the same verified catalog the quiz uses, and the URL updates as you switch models, so you can share the exact matchup you are looking at.

Can I compare three LLMs at once?

Yes — click "Add a third model" in the compare panel and the table becomes a three-way comparison. Every metric row then highlights the leader across all three, and each model gets its own "choose this one if…" column reasoned against the other two rather than against a single rival. Three is the maximum, because a fourth column stops being readable on a laptop screen and the useful question is almost always a shortlist of three. Three-model comparisons are shareable too: the URL becomes ?compare=model-a-vs-model-b-vs-model-c.

Where do I find one model's pricing, context window, and benchmarks?

Click any model name on this page and its full spec sheet opens in a pop-up — in the compare panel, in the comparison table headers, on the "best LLM for X" cards, and in the quiz results. The spec sheet shows full pricing (including cached input and batch discounts), context window and max output, throughput and time to first token, capability flags, deployment and compliance options, monthly cost at 10k / 100k / 1M calls, the trade-offs we know about, and which use cases we consider it the editorial pick for. Every spec sheet carries its own "verified" date.

Which LLM is best for coding in 2026?

As of July 2026, Claude Opus 5 is the strongest coding LLM — it tops SWE-bench Verified at 96.0% — but at $5 / $25 per 1M tokens most teams should reach for Claude Sonnet 5 ($3 / $15), which leads the price/performance tier on real-world coding and agentic benchmarks. GPT-5.6 Terra is a close third and got 20% cheaper on 2026-07-30 ($2 / $12) with strong tool use and fine-tuning support. For IDE-style fill-in-the-middle and autocomplete, Codestral 25.08 ($0.30 / $0.90 per 1M) is the best price/performance option with a 256K context window for large codebases.

Which LLM is best for RAG?

Gemini 3.1 Pro leads for RAG in 2026 because of its 2M-token context window — the largest available from any frontier lab — and strong needle-in-haystack recall. Claude Sonnet 5 is a strong second with 1M context, prompt caching (up to 90% cost cut on repeat prompts), and excellent grounding. For on-prem or open-weights RAG deployments, GLM-5.2 and DeepSeek V4 are the leading self-hostable options, both with 1M-token context windows.

Which LLM is best for AI agents and tool use?

Claude Sonnet 5 is the best LLM for autonomous agents and tool use in 2026, leading τ-bench and real-world multi-step tool-calling benchmarks. Use Claude Opus 5 for the hardest agent loops where reasoning depth justifies the cost; use GPT-5.6 Terra when you need OpenAI's tool ecosystem (Code Interpreter, File Search, Built-in Retrieval).

Which LLM is cheapest for production in 2026?

GPT-5.6 Luna at $0.20 / $1.20 per 1M tokens is the cheapest credible frontier-lab LLM as of 2026-07-30, when OpenAI cut its price by 80% — it undercuts Gemini 3.1 Flash-Lite ($0.25 / $1.50), the previous holder, with Claude Haiku 4.5 ($1 / $5) a tier above. At high volume, prompt caching matters more than list price — it cuts repeated-prompt cost by 80–90% and is available on Claude, GPT-5, and Gemini. For open-weights / self-hosted workloads, DeepSeek V4 Flash is cheapest at $0.14 / $0.28 per 1M (just $0.0028 per 1M for cached input) when China data-routing is acceptable; GLM-5.2 (Z.ai) is a strong open-weights alternative you can self-host to avoid third-party data routing.

Which LLM is best for maps and geospatial applications?

Gemini 3.1 Pro is the best LLM for maps and geospatial apps because Google's knowledge graph and Maps-adjacent training data give it a structural advantage in place names, routing, and travel reasoning. Gemini 3 Flash ($0.50 / $3) gives you the same geo advantage at production cost, and Gemini 3.6 Flash sits between them. Claude Sonnet 5 is a solid fallback when you need stronger tool use to chain Google Maps Platform APIs.

What is the best open-weights LLM in 2026?

GLM-5.2 (Z.ai) is the best general-purpose open-weights LLM as of July 2026 — a usable 1M-token context window and the strongest open-weights coding scores (top Terminal-Bench 2.1 and SWE-bench Pro, #1 on Design Arena's code leaderboard). DeepSeek V4 (Flash at $0.14 / $0.28, Pro at $0.435 / $0.87 per 1M tokens) is the leading open-weights reasoning family when China data-routing is acceptable. Codestral 25.08 remains the best open-weights coding specialist with a 256K context. Note that Alibaba's flagship Qwen3.7-Max is now closed-weight, so it no longer counts as an open-weights option. Meta's Llama 4 family is excluded from this picker because it underperforms current open-weights leaders on independent benchmarks.

Do you use an LLM to pick the LLM?

No — the picker is fully deterministic. It combines hard filters (modality, context window, compliance), soft scoring (benchmarks weighted by use case), and a curated editorial layer per use case. Every recommendation shows the full score breakdown. No inference calls are made, which means your answers stay on your device and the same inputs always produce the same ranking.

How fresh is the LLM catalog?

The catalog was last audited 2026-08-16 and every model's spec sheet carries its own "verified" date. New frontier models are added within a week of release; pricing and capability claims are reverified monthly against each provider's own pricing page, never an aggregator. If a model's verification date is more than 90 days old, treat pricing and capabilities as stale and confirm on the provider's page.

Why does the picker list older models like Claude Opus 4.8 or GPT-5.4?

Because they still answer API calls, and if you are running one in production you need to compare it against its replacement before migrating. Models superseded by a newer sibling are marked Legacy rather than deleted: they stay selectable in the compare panel and keep full spec sheets, but they are excluded from the quiz results and the "best LLM for X" rankings, so the picker never steers a new build onto one. A model is only removed from the picker when its provider actually retires it. Where a provider has published a retirement date, the spec sheet shows it.

How are the monthly cost estimates calculated?

Cost estimates combine four inputs: (1) your selected monthly call volume, (2) typical token shape per use case — for example RAG averages 8k input / 600 output tokens per call, coding averages 3k / 800, (3) published provider pricing per 1M tokens, and (4) realistic prompt-caching assumptions for models that support it. Batch-eligible workloads apply the provider's 50% batch discount when you select "batch is fine" for latency. The spec-sheet tiles use a flat 2k input / 600 output shape with caching excluded so models stay comparable. Figures are ballpark, not contractual — always run a pilot before committing to production.

Why does the picker sometimes override the benchmark winner?

Editorial picks override the benchmark leader when a model has provider-specific strengths that generic benchmarks miss. Gemini 3.1 Pro wins for maps because of Google's knowledge graph; Codestral 25.08 wins for IDE autocomplete because of its fill-in-the-middle training. Editorial boosts are explicitly labeled in the result card with the reason shown, so you can always see when an opinion is being applied and decide if you agree.