Technical Decision

Picking an LLM API for production

How to choose between Anthropic, OpenAI, Google Gemini and open-weight models: evals, October 2026 prices, limits, privacy, clouds and lock-in.

By Lance King · · 10 min read

A taster with a spoon and a clipboard of checkmarks stands beside a counter where four smiling robot chefs, each a different color, hold out bowls to be tasted.

This is for engineers and technical founders about to put a large language model behind a real feature: a support assistant, a document extractor, a coding agent, a classifier. The big three hosted APIs are Anthropic (Claude), OpenAI and Google (Gemini), and open-weight models (models whose weights you can download and run, served by a host or by you) are a fourth option. All of them are good. None is best at everything. The decision is which one does your task well enough, at a price and speed you can afford, under data terms your customers accept, without welding your code to it.

The short answer

Choose whichever model wins your own eval at a cost and latency you can live with. That beats any leaderboard.

Choose the provider your cloud already offers if procurement, data residency or a committed spend agreement matters more than the last few points of quality.

Choose a small model first for high-volume, narrow tasks like classification or extraction, and move up only where your eval shows you need to.

Choose an open-weight model if you need to control where the model runs, pin it forever, or fine-tune it, and you can accept the hosting work.

What actually matters

  1. Quality on your task. Public benchmarks measure someone else’s task. Build a small eval (more below).
  2. Price per million tokens. A token is roughly a word fragment. You pay separately for input and output, and output costs several times more, often five times.
  3. Latency. Time to first token and total time. Bigger models and more “thinking” are slower.
  4. Context window. How much text fits in one request. Most current models take about a million tokens, but some charge more past a threshold.
  5. Rate limits. How many requests and tokens per minute you get, and how fast that grows.
  6. Data retention and privacy. Whether prompts are stored, for how long, and whether they’re used for training.
  7. Cloud availability, features and lock-in. Whether you can buy it through AWS, Google Cloud or Azure, whether it does structured outputs and tool use well, and how hard it is to switch.

Prices and specs, side by side

Prices are standard per-million-token rates (input / output) from each provider’s official pricing page, as of October 2026. All three hosted providers offer 50% off for batch (asynchronous) jobs, and discounted rates for cached input.

Provider Model Input / output per 1M tokens Context window Notes
Anthropic Claude Fable 5.1 $10 / $50 1M Top tier; slowest in lineup
Anthropic Claude Opus 5.5 $4 / $20 1M Anthropic’s suggested default
Anthropic Claude Sonnet 5.5 $2 / $10 1M “Fast” in Anthropic’s latency labels
Anthropic Claude Haiku 4.5 $1 / $5 200K Fastest; retirement not sooner than Oct 15, 2026
OpenAI GPT-6 Astra $10 / $50 1.05M Above 272K input tokens: $20 / $75
OpenAI GPT-6.1 Sol $2 / $10 1.05M Above 272K: $4 / $15
OpenAI GPT-6 Luna $0.10 / $0.50 1.05M Above 272K: $0.20 / $0.75
Google Gemini 3.1 Pro (preview) $2 / $12 1,048,576 Above 200K: $4 / $18
Google Gemini 3.8 Flash $0.75 / $3.75 1,048,576 Rises to $1.50 / $7.50 on Jan 1, 2027
Google Gemini 3.5 Flash-Lite $0.30 / $2.50 See model page Small, low-cost tier
Open-weight gpt-oss-120B on Together AI $0.15 / $0.60 Varies by host One example; prices vary by host

A few things jump out. The flagships cluster at similar prices, so price rarely decides between them. The small models differ by an order of magnitude, so price often decides between them. Gemini 3.1 Pro is still labeled preview, which matters if your company won’t ship on preview models. Gemini 3.8 Flash has an introductory price that doubles in January, so budget for the later number. And Anthropic’s model page lists Haiku 4.5’s retirement as “not sooner than October 15, 2026,” so check the deprecation schedule before building on it.

Regional processing costs extra almost everywhere. OpenAI charges a 10% uplift on data-residency endpoints for models released on or after March 5, 2026. Anthropic charges 1.1x for US-only inference. On Google Cloud, regional and multi-region endpoints for Claude carry a 10% premium.

Quality: build a small eval

An eval is a fixed set of inputs with a way to score outputs. It doesn’t need to be fancy:

  • Collect 50 to 200 real examples of the task, including the ugly ones: long inputs, missing fields, hostile users.
  • Write down what a good answer is. For extraction, that’s exact fields. For writing, a short rubric.
  • Score automatically where you can (exact match, JSON validity, tests passing). Where you can’t, use human review or a grader model checked against human labels.
  • Run every candidate model through the same set and record quality, cost per task and latency together.

Keep the eval in your repo next to the prompts. It’s what lets you switch models later without guessing, and it catches regressions when a provider updates a model.

Latency

Latency depends so much on your own prompts that published figures are only a starting point. Anthropic, for example, labels its models from “Fastest” (Haiku) to “Slower” (Fable) and notes that real latency depends on prompt length, output length and thinking effort. Measure it yourself: p50 and p95 time to first token and total time, from the region your servers run in, with your real prompt sizes. Streaming hides a lot of latency in chat interfaces. It hides none in a background job that waits for the full answer.

Context windows

A million tokens of context is now normal at all three. Two cautions. First, price changes past a threshold: OpenAI’s GPT-6 models switch to higher “long context” rates above 272K input tokens, and Gemini 3.1 Pro above 200K. Second, a huge context isn’t free retrieval. Sending a whole knowledge base in every request costs money and time. Caching repeated input helps: Anthropic says cache reads on most models don’t count toward its input rate limit, and all three discount cached tokens.

Rate limits

Prices in this section are as of October 2026.

All three use tiers that rise with spend and history:

  • OpenAI has a free tier and Tiers 1 to 5, reached at $5, $50, $100, $250 and $1,000 paid. Limits are measured in requests and tokens per minute and per day.
  • Anthropic places organizations on Start, Build and Scale tiers, with monthly spend caps of $500, $1,000 and $200,000, and new accounts may begin on a lower Evaluation tier. Limits are per model, in requests, input tokens and output tokens per minute.
  • Google has a free tier and Tiers 1 to 3. Tier 2 needs $100 spent and 3 days since first payment; Tier 3 needs $1,000 and 30 days.

The practical lesson: a launch can hit limits on day one. Spend enough ahead of time to reach a comfortable tier, request increases early, and handle 429 (rate limited) errors with backoff.

Data retention and privacy

Read the terms for the exact product you use, because the same model can come with different terms on different platforms.

  • OpenAI says API data isn’t used for training unless you opt in. Abuse-monitoring logs are kept up to 30 days. Zero Data Retention is available to eligible customers with approval.
  • Anthropic says retained data is never used for training without your express permission, and its commercial policy deletes API inputs and outputs within 30 days. Zero data retention is available for eligible features, except on its Fable and Mythos models, which require 30-day retention.
  • Google treats free and paid use differently. On unpaid services, prompts and responses may be used to improve Google products and read by human reviewers. On paid services, Google says it doesn’t use them to improve products and logs them for a limited period for abuse detection. Don’t send customer data through a free-tier key.
  • Through a cloud, the cloud provider is usually the data processor. Anthropic’s docs say that on Amazon Bedrock and Google Cloud, those platforms’ data terms apply.

Cloud platforms

If your company already buys from one cloud, buying the model there can mean one bill, existing compliance paperwork and private networking.

  • Claude is on Amazon Bedrock, Google Cloud’s Gemini Enterprise Agent Platform, and Microsoft Foundry, plus Anthropic’s own API.
  • OpenAI’s GPT-6 models are on Microsoft Foundry, and Bedrock lists GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna.
  • Gemini is available through Google’s Gemini API and Google Cloud. Bedrock’s model list includes Google’s open Gemma models but not Gemini.
  • Open-weight models such as gpt-oss, Llama, Qwen, DeepSeek and Mistral are on Bedrock and on specialist hosts like Together AI.

Features aren’t always identical across platforms. For example, Anthropic’s docs list the Message Batches API and Files API as unsupported for Claude on Google Cloud. Check the feature table for your platform, not just the model.

Structured outputs and tool use

If your code parses model output, you want guaranteed structure. All three now offer it. OpenAI’s Structured Outputs “ensures the model will always generate responses that adhere to your supplied JSON Schema,” with a strict mode for function calling. Anthropic’s structured outputs use constrained decoding for JSON output and strict tool use, with documented limits on schema features. Gemini 3.8 Flash and 3.1 Pro list structured outputs and function calling as supported. Test your actual schemas: limits on things like optional fields, unions and regex patterns differ.

Which one for you

Solo founder shipping a feature. Pick two candidates, run your eval, ship the cheaper one that passes. Use the provider’s own API for speed of setup, but keep a paid key from day one.

Growing startup with real volume. Split tasks by difficulty. Small models handle routing, classification and extraction. A mid-tier model handles most generation. A flagship only for the hard cases. Move background work to batch for the 50% discount.

Regulated company. Start from where your data is allowed to go. Buying through your existing cloud often settles retention, residency and contracts faster than a new vendor review. Confirm whether the model you want supports zero retention or regional processing on that platform, and price in the regional uplift.

Team that needs control. Open-weight models let you choose the host, pin a version indefinitely and fine-tune. You give up some quality at the top end and take on capacity planning. See build vs. buy before you decide to host it yourself.

Mistakes to avoid

  • Choosing from a leaderboard. Your eval is the only benchmark that matters.
  • Comparing input prices only. Output tokens usually dominate cost, and reasoning models can produce many of them.
  • Ignoring preview labels and retirement dates. A model that changes or disappears mid-quarter is a production incident.
  • Testing on the free tier, launching on paid, without noticing the data terms changed.
  • Leaving rate limits until launch day.

Avoiding lock-in

Lock-in with LLM APIs is mostly self-inflicted. Some habits that keep you portable:

  • Write a thin interface of your own. One function like complete(task, input) -> result that your app calls, with one adapter per provider behind it. Keep it small; you only need the features you use.
  • Don’t rely on compatibility shims for production. Google’s OpenAI-compatible endpoint is in beta with limited features. Anthropic says its OpenAI SDK compatibility is “primarily intended to test and compare model capabilities” and ignores strict on tools and doesn’t support prompt caching. They’re great for a quick eval run, not as your main path.
  • Keep prompts and evals in your repo, versioned, not in a vendor console. Expect to retune prompts per model; that’s normal.
  • Store raw inputs and outputs (where your privacy terms allow) so you can replay them against a new model.
  • Record the choice in a technical decision record with the eval scores and prices, so you can revisit it when the next model ships.

A quick checklist

  • Build a 50 to 200 example eval from real data.
  • Run at least two providers and one small model through it.
  • Calculate cost per task, including output and long-context rates.
  • Measure p50 and p95 latency from your own region.
  • Check rate-limit tiers and request increases before launch.
  • Confirm retention, training and residency terms for the exact platform.
  • Check preview status and retirement dates.
  • Put a thin interface in front of the API and keep prompts in git.
  • Write down the decision.

Sources