Which model should you use?

Start with Sonnet for general work, Luna for inexpensive repetition, and gpt-oss when the data must stay with you. Use a managed API unless your constraints justify running the model yourself.

The Tandem shortlist

Start here. Change when the work says to.

These are our starting picks, not a benchmark ranking. Match your work below, then apply the privacy rules before sending data.

Anthropic logo Managed API

An assistant for your team

Small product team · mixed knowledge work

Start withClaude Sonnet 5

Our default for questions over company documents, drafting, and everyday analysis. Start here before paying for a flagship on every request.

The trade-off. Try Luna for repetitive work. Fix retrieval and document permissions before blaming the model for a bad answer.

OpenAI logo Managed API

High-volume customer support

Lean team · tight per-conversation budget

Start withGPT-5.6 Luna

Our first pick for bounded support and classification. Keep the routine path inexpensive, with a human escalation route.

The trade-off. Move difficult cases to Sonnet if it reduces errors in your evaluation. Don’t let either model improvise refunds or policy.

Google logo Managed API

Extracting data from documents

Operations team · volume over deep reasoning

Start withGemini 3.5 Flash-Lite

Start with Flash-Lite for repeatable extraction and document triage. We’d rather pay for validation than premium reasoning on every form.

The trade-off. Try Sonnet when layout or ambiguity drives costly errors. Use the paid API; its data terms differ from the free tier.

Anthropic logo Managed API

Engineering agents

Experienced reviewers · errors cost more than tokens

Start withClaude Opus 5

Our starting pick for difficult code changes. We use Claude in delivery; compare GPT-6 Astra on the same repository tasks before committing.

The trade-off. Use Sonnet for routine edits. Only pay the Astra or Fable premium when fewer failed attempts justify it.

OpenAI logo Self-hosted

Data cannot leave your boundary

Infrastructure team · 128–256 GB system to evaluate

Start withgpt-oss-120b

Start a private text-reasoning evaluation here. Self-host the full path, including retrieval and tools, rather than sending sensitive context to an API.

The trade-off. You own serving, patching, logs, and access control. Measure runtime memory and speed before buying hardware; weights alone prove neither.

OpenAI logo Self-hosted

A smaller offline pilot

Developer-led experiment · limited local memory

Start withgpt-oss-20b

Our first local candidate for narrow text tasks when the 120b model is too much. Keep the scope small enough to inspect its mistakes.

The trade-off. Move to 120b only if quality earns the extra memory. If nobody can operate a local service, use approved hosting or postpone sensitive-data use.

Already standardized on Azure, Bedrock, or Google Cloud? Favor an eligible model on your approved platform before adding another vendor. Confirm the exact model and regional endpoint; direct-API terms don’t transfer.

Worth a place in the evaluation

Where the other families belong

A shortlist needs challengers. Here’s when we’d bring these families in, and what keeps them from being our default.

GrokResearch involving X
Put Grok 4.6 with X Search on the list when current public conversation is part of the job. We wouldn’t add another vendor for ordinary document Q&A alone. Personal Data requires ZDR under xAI’s terms; check tool eligibility and charges separately.
LlamaAn AWS-first team
Try Llama 4 Scout on Bedrock against Luna for bounded text work if AWS is already your approved home. The low hosted rate makes that comparison worthwhile without operating GPUs. Self-hosting brings infrastructure work and Meta’s license conditions; it isn’t the same procurement decision.
QwenVolume where cost matters
Qwen3.8-Flash belongs in a high-volume extraction or classification evaluation. Its price is a reason to test it, not a reason to relax your controls. The Singapore/International offering allows cross-border inference and publishes no contractual ZDR option. Keep strict-retention workloads off that path.
DeepSeekNon-sensitive reasoning work
Try V4 Flash on public or synthetic inputs when you can measure reasoning quality against cost. We wouldn’t make its direct API the default for confidential work while training exclusions and contractual ZDR remain unconfirmed. An open checkpoint is another deployment option, with its own hardware and operating bill.
GLMAn inexpensive hosted challenger
Put GLM-5.3-Flash against Luna for non-sensitive support or extraction. Judge its undiscounted price, not a temporary promotion. We’d hold confidential workloads until Z.ai reconciles its no-content-storage promise with automatic caching. Cheap hosted tokens don’t establish cheap self-hosting.

Non-negotiables first

Privacy can change the answer.

“Not used for training” doesn’t mean “not retained.” Apply your strictest requirement first; a cheaper model doesn’t get an exception.

No outside processing

Choose local gpt-oss, or a private deployment your policy permits. Keep prompts, retrieval, tools, telemetry, and backups inside that boundary. No operating team? Don’t ship sensitive data yet.

Hosted, but no retention

Evaluate Sonnet or Luna only with approved zero data retention (ZDR) and eligible features. Check caches, files, tools, and safety/legal exceptions. Rule out Fable 5.1 unless Anthropic expressly grants a ZDR exception.

Regulated or region-bound

Choose the contract and deployment first, then an eligible model. PHI needs a signed BAA and covered features; PII and financial data have different obligations. Regional storage alone doesn’t guarantee regional processing.

The service behind the model

Enterprise/API text services, not consumer subscriptions. Policy links go directly to vendor sources.
Named service Default retention Zero data retention (ZDR) What changes the decision
Anthropic direct API

Unverified. Docs conflict: deletion within 30 days versus no conversation retention by default. Policy

Conflicting platform docs

By request, per organization, for eligible Messages and Token Counting features. Terms

Fable 5.1 requires 30-day retention unless Anthropic expressly grants an exception. Standard ZDR excludes Console, Files, Batch, and code execution.

Fable terms
OpenAI direct API

Abuse logs up to 30 days; Responses state at least 30 days by default. Policy

Prior approval required; model, feature, and safety exceptions remain. Terms

store=false does not remove abuse logs. Ineligible features can retain state; image/file review and notified safety-retention exceptions still apply.

Gemini Developer API, paid projects

Abuse-monitoring content retained 55 days; flagged content may receive human review. Policy

Project approval required; separately control grounding, storage, files, and caches. Terms

Not free Gemini or Vertex AI. Avoid Search/Maps grounding and Live session resumption; disable Interactions storage. Google permits 24-hour RAM caching under ZDR.

Grok business API

30 days by default; agreed periods and contract exceptions can change this. Policy

Required for Personal Data; team-wide where available, with stateful features excluded. Terms

No stateful Responses, Files, Collections, Batch, or deferred completions. Contract review must reconcile the DPA’s backup/storage language with the specific ZDR promise.

Llama 4 on Amazon Bedrock, US geo

No default model-input/output storage for this stateless text scope, with invocation logging off. Policy

Logging controls

Documented for these Llama 4 models in this scope. Terms

Uses bedrock-runtime from us-east-1 with a US geo profile. Excludes images, agents, and external tools; no blanket claim about other AWS services.

Qwen on Model Studio, Singapore / International

Audit metadata queryable for 30 days; content deletion deadline unspecified. Inference logging off. Policy

Not published. No contractual zero-content-retention option established in the checked documents. Terms

The metadata query window does not prove content deletion. Singapore storage and API access do not prevent cross-border inference under International scope.

Region scope
DeepSeek Open Platform API

API-wide retention limit unresolved. Automatic disk caches usually clear after hours to days. Policy

Unverified. An API-wide contractual option remains unconfirmed. Terms

Cache expiry does not establish payload/log deletion. API-wide default training exclusion is also unconfirmed; no negotiated enterprise addendum or selectable region is assumed.

Open Platform terms
GLM on Z.ai business API

DPA promises no content storage, but automatic cache retention remains unresolved. Policy

Cache documentation

Unverified. End-to-end zero retention remains unconfirmed, including cache state. Terms

Other customer data may persist. US-use terms prohibit HIPAA PHI and GLBA nonpublic personal information. Excludes consumer chat and the subscription Coding Plan.

Usage restrictions

ZDR is scoped, not a universal deletion guarantee. Check application logs, caches, tools, safety reviews, and legal exceptions. Your exact service, settings, and current contract terms control. And, if you self-host, storage and deletion are your responsibility.

The practical trade-offs

What you pay. What you have to run.

Compare cost per accepted result, including retries and review. Local weights have no token bill, but hardware and the people operating it do.

Hosted API prices

Selected hosted API rates, USD per 1M tokens. Standard synchronous text, uncached input. Not a cheapest-provider ranking or rate guarantee.
Input Output One linear scale: $0–$50
  1. Qwen3.8-Flash (qwen3.8-flash) Alibaba Cloud · Direct vendor API
    Input $0.15 Output $0.47
  2. GLM-5.3-Flash Zhipu / Z.ai · Direct vendor API
    Input $0.15 Output $0.5
  3. Llama 4 Scout 17B Instruct Meta · AWS Bedrock
    Input $0.17 Output $0.66
  4. GPT-5.6 Luna OpenAI · Direct vendor API
    Input $0.2 Output $1.2
  5. Gemini 3.5 Flash-Lite Google · Direct vendor API
    Input $0.3 Output $2.5
  6. DeepSeek-V4-Flash-0731 DeepSeek · Direct vendor API
    Input $0.44 Output $1.32
  7. Grok 4.3 (grok-4.3) xAI · Direct vendor API
    Input $1.25 Output $2.5
  8. Claude Sonnet 5 Anthropic · Direct vendor API
    Input $2 Output $10
  9. Claude Opus 5 Anthropic · Direct vendor API
    Input $5 Output $25
  10. GPT-6 Astra OpenAI · Direct vendor API
    Input $10 Output $50

Ordered by input price, not quality. Rate scopes and exceptions.

Weights only

0 128 GB 256 GB
  1. gpt-oss-20b 13.8 GB
    Published native payload · source
  2. Gemma 4 31B IT 15.4 GB
    Nominal 4-bit estimate · source
  3. Llama 4 Scout 17B Instruct 54.5 GB
    Nominal 4-bit estimate · source
  4. Qwen3.8-27B 55.6 GB
    Published native payload · source
  5. gpt-oss-120b 65.2 GB
    Published native payload · source
  6. Llama 4 Maverick 17B Instruct 200.0 GB
    Nominal 4-bit estimate · source
Weights in decimal GB exclude OS, runtime and KV cache. Nominal 4-bit estimates count all parameters (all experts), not files, and exclude quantization metadata and higher-precision tensors. Native payloads exclude file headers. Fit and throughput are unmeasured; multiple machines don’t automatically pool memory.

Before you commit

Make the model earn its place.

Run the same real tasks through your starting pick and one alternative. Compare errors, latency, and total cost. Agents with production write access need least privilege, approvals, audit logs, and a recovery path regardless of model.

Get help choosing and deploying AI →