The Tandem shortlist
Start here. Change when the work says to.
These are our starting picks, not a benchmark ranking. Match your work below, then apply the privacy rules before sending data.
An assistant for your team
Small product team · mixed knowledge work
Start withClaude Sonnet 5
Our default for questions over company documents, drafting, and everyday analysis. Start here before paying for a flagship on every request.
The trade-off. Try Luna for repetitive work. Fix retrieval and document permissions before blaming the model for a bad answer.
High-volume customer support
Lean team · tight per-conversation budget
Start withGPT-5.6 Luna
Our first pick for bounded support and classification. Keep the routine path inexpensive, with a human escalation route.
The trade-off. Move difficult cases to Sonnet if it reduces errors in your evaluation. Don’t let either model improvise refunds or policy.
Extracting data from documents
Operations team · volume over deep reasoning
Start withGemini 3.5 Flash-Lite
Start with Flash-Lite for repeatable extraction and document triage. We’d rather pay for validation than premium reasoning on every form.
The trade-off. Try Sonnet when layout or ambiguity drives costly errors. Use the paid API; its data terms differ from the free tier.
Engineering agents
Experienced reviewers · errors cost more than tokens
Start withClaude Opus 5
Our starting pick for difficult code changes. We use Claude in delivery; compare GPT-6 Astra on the same repository tasks before committing.
The trade-off. Use Sonnet for routine edits. Only pay the Astra or Fable premium when fewer failed attempts justify it.
Data cannot leave your boundary
Infrastructure team · 128–256 GB system to evaluate
Start withgpt-oss-120b
Start a private text-reasoning evaluation here. Self-host the full path, including retrieval and tools, rather than sending sensitive context to an API.
The trade-off. You own serving, patching, logs, and access control. Measure runtime memory and speed before buying hardware; weights alone prove neither.
A smaller offline pilot
Developer-led experiment · limited local memory
Start withgpt-oss-20b
Our first local candidate for narrow text tasks when the 120b model is too much. Keep the scope small enough to inspect its mistakes.
The trade-off. Move to 120b only if quality earns the extra memory. If nobody can operate a local service, use approved hosting or postpone sensitive-data use.
Already standardized on Azure, Bedrock, or Google Cloud? Favor an eligible model on your approved platform before adding another vendor. Confirm the exact model and regional endpoint; direct-API terms don’t transfer.
Worth a place in the evaluation
Where the other families belong
A shortlist needs challengers. Here’s when we’d bring these families in, and what keeps them from being our default.
- GrokResearch involving X
- Put Grok 4.6 with X Search on the list when current public conversation is part of the job. We wouldn’t add another vendor for ordinary document Q&A alone. Personal Data requires ZDR under xAI’s terms; check tool eligibility and charges separately.
- LlamaAn AWS-first team
- Try Llama 4 Scout on Bedrock against Luna for bounded text work if AWS is already your approved home. The low hosted rate makes that comparison worthwhile without operating GPUs. Self-hosting brings infrastructure work and Meta’s license conditions; it isn’t the same procurement decision.
- QwenVolume where cost matters
- Qwen3.8-Flash belongs in a high-volume extraction or classification evaluation. Its price is a reason to test it, not a reason to relax your controls. The Singapore/International offering allows cross-border inference and publishes no contractual ZDR option. Keep strict-retention workloads off that path.
- DeepSeekNon-sensitive reasoning work
- Try V4 Flash on public or synthetic inputs when you can measure reasoning quality against cost. We wouldn’t make its direct API the default for confidential work while training exclusions and contractual ZDR remain unconfirmed. An open checkpoint is another deployment option, with its own hardware and operating bill.
- GLMAn inexpensive hosted challenger
- Put GLM-5.3-Flash against Luna for non-sensitive support or extraction. Judge its undiscounted price, not a temporary promotion. We’d hold confidential workloads until Z.ai reconciles its no-content-storage promise with automatic caching. Cheap hosted tokens don’t establish cheap self-hosting.
Non-negotiables first
Privacy can change the answer.
“Not used for training” doesn’t mean “not retained.” Apply your strictest requirement first; a cheaper model doesn’t get an exception.
No outside processing
Choose local gpt-oss, or a private deployment your policy permits. Keep prompts, retrieval, tools, telemetry, and backups inside that boundary. No operating team? Don’t ship sensitive data yet.
Hosted, but no retention
Evaluate Sonnet or Luna only with approved zero data retention (ZDR) and eligible features. Check caches, files, tools, and safety/legal exceptions. Rule out Fable 5.1 unless Anthropic expressly grants a ZDR exception.
Regulated or region-bound
Choose the contract and deployment first, then an eligible model. PHI needs a signed BAA and covered features; PII and financial data have different obligations. Regional storage alone doesn’t guarantee regional processing.
The service behind the model
| Named service | Default retention | Zero data retention (ZDR) | What changes the decision |
|---|---|---|---|
| Anthropic direct API | Unverified. Docs conflict: deletion within 30 days versus no conversation retention by default. Policy Conflicting platform docs | By request, per organization, for eligible Messages and Token Counting features. Terms | Fable 5.1 requires 30-day retention unless Anthropic expressly grants an exception. Standard ZDR excludes Console, Files, Batch, and code execution. Fable terms |
| OpenAI direct API | Abuse logs up to 30 days; Responses state at least 30 days by default. Policy | Prior approval required; model, feature, and safety exceptions remain. Terms | store=false does not remove abuse logs. Ineligible features can retain state; image/file review and notified safety-retention exceptions still apply. |
| Gemini Developer API, paid projects | Abuse-monitoring content retained 55 days; flagged content may receive human review. Policy | Project approval required; separately control grounding, storage, files, and caches. Terms | Not free Gemini or Vertex AI. Avoid Search/Maps grounding and Live session resumption; disable Interactions storage. Google permits 24-hour RAM caching under ZDR. |
| Grok business API | 30 days by default; agreed periods and contract exceptions can change this. Policy | Required for Personal Data; team-wide where available, with stateful features excluded. Terms | No stateful Responses, Files, Collections, Batch, or deferred completions. Contract review must reconcile the DPA’s backup/storage language with the specific ZDR promise. |
| Llama 4 on Amazon Bedrock, US geo | No default model-input/output storage for this stateless text scope, with invocation logging off. Policy Logging controls | Documented for these Llama 4 models in this scope. Terms | Uses bedrock-runtime from us-east-1 with a US geo profile. Excludes images, agents, and external tools; no blanket claim about other AWS services. |
| Qwen on Model Studio, Singapore / International | Audit metadata queryable for 30 days; content deletion deadline unspecified. Inference logging off. Policy | Not published. No contractual zero-content-retention option established in the checked documents. Terms | The metadata query window does not prove content deletion. Singapore storage and API access do not prevent cross-border inference under International scope. Region scope |
| DeepSeek Open Platform API | API-wide retention limit unresolved. Automatic disk caches usually clear after hours to days. Policy | Unverified. An API-wide contractual option remains unconfirmed. Terms | Cache expiry does not establish payload/log deletion. API-wide default training exclusion is also unconfirmed; no negotiated enterprise addendum or selectable region is assumed. Open Platform terms |
| GLM on Z.ai business API | DPA promises no content storage, but automatic cache retention remains unresolved. Policy Cache documentation | Unverified. End-to-end zero retention remains unconfirmed, including cache state. Terms | Other customer data may persist. US-use terms prohibit HIPAA PHI and GLBA nonpublic personal information. Excludes consumer chat and the subscription Coding Plan. Usage restrictions |
ZDR is scoped, not a universal deletion guarantee. Check application logs, caches, tools, safety reviews, and legal exceptions. Your exact service, settings, and current contract terms control. And, if you self-host, storage and deletion are your responsibility.
The practical trade-offs
What you pay. What you have to run.
Compare cost per accepted result, including retries and review. Local weights have no token bill, but hardware and the people operating it do.
Hosted API prices
- Qwen3.8-Flash (qwen3.8-flash) Alibaba Cloud · Direct vendor APIInput $0.15 Output $0.47
- GLM-5.3-Flash Zhipu / Z.ai · Direct vendor APIInput $0.15 Output $0.5
- Llama 4 Scout 17B Instruct Meta · AWS BedrockInput $0.17 Output $0.66
- GPT-5.6 Luna OpenAI · Direct vendor APIInput $0.2 Output $1.2
- Gemini 3.5 Flash-Lite Google · Direct vendor APIInput $0.3 Output $2.5
- DeepSeek-V4-Flash-0731 DeepSeek · Direct vendor APIInput $0.44 Output $1.32
- Grok 4.3 (grok-4.3) xAI · Direct vendor APIInput $1.25 Output $2.5
- Claude Sonnet 5 Anthropic · Direct vendor APIInput $2 Output $10
- Claude Opus 5 Anthropic · Direct vendor APIInput $5 Output $25
- GPT-6 Astra OpenAI · Direct vendor APIInput $10 Output $50
Ordered by input price, not quality. Rate scopes and exceptions.
Weights only
- gpt-oss-20b 13.8 GBPublished native payload · source
- Gemma 4 31B IT 15.4 GBNominal 4-bit estimate · source
- Llama 4 Scout 17B Instruct 54.5 GBNominal 4-bit estimate · source
- Qwen3.8-27B 55.6 GBPublished native payload · source
- gpt-oss-120b 65.2 GBPublished native payload · source
- Llama 4 Maverick 17B Instruct 200.0 GBNominal 4-bit estimate · source
Before you commit
Make the model earn its place.
Run the same real tasks through your starting pick and one alternative. Compare errors, latency, and total cost. Agents with production write access need least privilege, approvals, audit logs, and a recovery path regardless of model.
Get help choosing and deploying AI →