The Tandem shortlist
Start here. Change when the work says to.
Starting picks, not a benchmark ranking. Match your work, then check the privacy rules before you send anything real. Every pick has a row in the matrix below.
An assistant for your team
Small product team · mixed knowledge work
Start withClaude Sonnet 5
Our default for questions over company documents, drafting, and everyday analysis. Don’t pay flagship rates on every request.
The trade-off. Try Luna for the repetitive work. And, fix retrieval and document permissions before blaming the model for a bad answer.
High-volume customer support
Lean team · tight per-conversation budget
Start withGPT-5.6 Luna
Our first pick for bounded support and classification. Keep the routine path cheap, with a human escalation route.
The trade-off. Send hard cases to Sonnet if your evaluation shows fewer errors. Neither model gets to improvise refunds or policy.
Extracting data from documents
Operations team · volume over deep reasoning
Start withGemini 3.5 Flash-Lite
Start here for repeatable extraction and document triage. We’d rather pay for validation than for premium reasoning on every form.
The trade-off. Move to Sonnet when layout or ambiguity drives costly errors. Use the paid API; the free tier has different data terms.
Engineering agents
Experienced reviewers · errors cost more than tokens
Start withClaude Opus 5
Our starting pick for difficult code changes. We use Claude in our own delivery, so run GPT-6 Astra on the same repository tasks before you commit.
The trade-off. Use Sonnet for routine edits. Astra and Fable 5.1 list at twice the price of Opus; pay that only when fewer failed attempts earn it back.
Data cannot leave your boundary
Infrastructure team · a GPU server to evaluate on
Start withgpt-oss-120b
Start a private text-reasoning evaluation here. Self-host the whole path, retrieval and tools included, instead of sending sensitive context to an API.
The trade-off. You own serving, patching, logs, and access control. Measure memory and speed on real hardware before buying any; the weights are the floor, not the budget.
A smaller offline pilot
Developer-led experiment · limited local memory
Start withgpt-oss-20b
Our first local candidate for narrow text tasks when 120b is too much machine. Keep the scope small enough to inspect every mistake.
The trade-off. Move up to 120b only if quality earns the memory. If nobody can operate a local service, use approved hosting or wait.
Already on Azure, Bedrock, or Google Cloud? Favor an eligible model on your approved platform before adding a vendor. Direct-API terms don’t transfer, so confirm the exact model and regional endpoint.
Non-negotiables first
Privacy can change the answer.
“Not used for training” doesn’t mean “not retained.” Apply your strictest requirement first. A cheaper model doesn’t get an exception.
Nothing leaves your boundary
Run gpt-oss locally, or a private deployment your policy permits. Keep prompts, retrieval, tools, telemetry, and backups inside it. No team to operate that? Don’t ship sensitive data yet.
Hosted, but nothing kept
Use Sonnet or Luna only with approved zero data retention (ZDR) and eligible features. Check caches, files, tools, and the safety and legal exceptions. Fable 5.1 is out unless Anthropic grants a written exception.
Regulated or region-bound
Pick the contract and deployment first, then an eligible model. PHI needs a signed BAA and covered features. Regional storage alone doesn’t guarantee regional processing.
The matrix
What each model costs, and what happens to your data.
Every model named on this page, with the terms that decide most enterprise deals. Read across a row for one model; read down the zero-retention and BAA columns to shorten your list fast.
| Model | Price in / out7 | Trains on your data | Kept by default | Zero retention | HIPAA BAA | Best for |
|---|---|---|---|---|---|---|
| Managed API | ||||||
| Claude Sonnet 5 Shortlist pick | $2 / $10 | No | Up to 30 days 1 | By request | Available | Everyday knowledge work |
| Claude Opus 5 Shortlist pick | $5 / $25 | No | Up to 30 days 1 | By request | Available | Hard code changes |
| Claude Fable 5.1 | $10 / $50 | No | 30 days, required 2 | Only by exception 2 | Available | When fewer retries pay for double the price |
| GPT-5.6 Luna Shortlist pick | $0.20 / $1.20 | No | Up to 30 days | By approval | Available | Cheap, bounded support |
| GPT-6 Astra | $10 / $50 | No | Up to 30 days | By approval | Available | Run it against Opus on your own repo |
| Gemini 3.5 Flash-Lite Shortlist pick | $0.30 / $2.50 | No 3 | 55 days | By project approval 3 | Not confirmed 3 | Document extraction at volume |
| Grok 4.6 | $2 / $6 | No | 30 days | Required for personal data 4 | With ZDR | Research that needs X |
| Llama 4 Scout 17B Instruct | $0.17 / $0.66 | No | None, logging off 5 | Default | Via AWS | AWS-first teams |
| Qwen3.8-Flash | $0.15 / $0.47 | No | No deadline stated 6 | None published | Not for regulated use | Cost-first volume, nothing sensitive |
| DeepSeek-V4-Flash-0731 | $0.44 / $1.32 | Not confirmed 6 | No limit stated | Not confirmed | Not confirmed | Reasoning on public inputs |
| GLM-5.3-Flash | $0.15 / $0.50 | No | None promised, cache open 6 | Not confirmed | PHI prohibited 6 | Cheap hosted challenger, nothing sensitive |
| Self-hosted | ||||||
| gpt-oss-120b Shortlist pick | No token bill. 65.2 GB of weights, before runtime and context memory. | Never sent | You control it | You run it | Your controls | Data that cannot leave |
| gpt-oss-20b Shortlist pick | No token bill. 13.8 GB of weights, before runtime and context memory. | Never sent | You control it | You run it | Your controls | A small offline pilot |
- Anthropic’s two policy pages describe standard retention differently. One says content isn’t retained by default, the other says it’s deleted within 30 days. We show the conservative reading; flagged content can be held longer.
- Fable 5.1 must keep inputs and outputs for 30 days and isn’t offered under zero data retention without a written exception from Anthropic. Flagged content can be kept up to 2 years.
- Paid Gemini Developer API only; the free tier and Vertex AI carry different terms. Abuse-monitoring data may train Google’s policy-enforcement models. Under ZDR, avoid Search and Maps grounding and Live session resumption; Google keeps a 24-hour in-memory cache.
- xAI requires ZDR before you send personal data. Stateful features such as Files, Batch, and Collections sit outside it.
- Bedrock in US East with a US geo profile, invocation logging off, text only. Meta and Alibaba also publish open weights for Llama and Qwen; self-hosting them is a separate procurement decision with its own license terms.
- Qwen International stores data in Singapore but allows cross-border inference. Z.ai’s DPA promises no content storage while its automatic cache retention stays unresolved, and its US terms prohibit PHI and GLBA data. DeepSeek processes personal data in China; an API-wide training exclusion and a zero-retention option weren’t confirmed in the documents we checked.
- List rates for standard synchronous text with uncached input, excluding batch, caching, long-context bands, and promotions. DeepSeek shows peak rates; off-peak is half. GLM shows undiscounted rates. Gemini and Qwen output includes thinking tokens.
Before you commit
Make the model earn its place.
Run the same real tasks through your starting pick and one alternative. Compare errors, latency, and total cost, retries and review included. Any agent with production write access needs least privilege, approvals, audit logs, and a way back, whatever the model.
Get help choosing and deploying AI