Batch API processing: when the 50% discount is worth it
The batch discount is the only line on an LLM bill you can halve without changing the model, the prompt, or the output. Prompt trimming costs you context. Dropping a tier costs you quality, and you have to prove it did not. Batch costs you time, and only on work where nobody is watching the clock. That asymmetry makes it the first saving to reach for and the one most teams leave on the table, because nobody has ever sat down and sorted their features into the two piles.
So the real question is never whether batch is worth it — at −50% it plainly is — but how much of your traffic can tolerate the wait. That is an audit, not a decision. Here is one, run end to end against a plausible product.
Six features, two piles
The product is a customer-feedback platform with roughly 900 business accounts. It has six AI features, built over two years by three different engineers, and they have never been looked at together.
The in-app chat assistant answers questions about an account's own ticket history. A human is staring at a spinner. It runs 220,000 messages a month at about 4,000 input tokens (the question plus conversation history plus retrieved context) and 350 output tokens, on Claude Sonnet 5. Interactive. Not a candidate, and never will be.
Inline reply suggestions fire as a support agent types, on Gemini 2.5 Flash-Lite, roughly 3.1 million requests a month at 600 in and 40 out. This is the most latency-sensitive thing in the product — a suggestion that arrives after the agent has finished typing is worse than no suggestion. Interactive.
The "summarise this thread" button produces a précis of a long ticket on demand: 60,000 uses a month, about 9,000 tokens in, 600 out, on Claude Sonnet 5. Someone clicked a button and is waiting. Interactive — though note it is the one feature on this list where a product decision could move it. If summaries were pre-generated overnight for every thread that changed that day, the button would become instant and cheap. Hold that thought.
Nightly ticket classification tags the previous day's tickets with sentiment and theme: 40,000 tickets a day, 1,200 tokens in, 150 out, on Claude Haiku 4.5. The results are read by humans in a dashboard the following morning. Nothing consumes them between 02:00 and 08:00. Deferrable.
Weekly executive summaries generate one report per account every Monday: 900 reports, about 25,000 input tokens each (a week of aggregated ticket data), 1,200 out, on Claude Sonnet 5. The job is kicked off on Sunday evening for Monday delivery. Deferrable.
The taxonomy backfill is the one nobody budgeted for. When the theme taxonomy changes, 2.4 million archived tickets have to be re-classified — same prompt shape as the nightly job, 1,200 in and 150 out on Claude Haiku 4.5. It runs perhaps twice a year and takes as long as it takes. Deferrable, emphatically.
Three interactive, three deferrable. That ratio is typical: the loud features are interactive, and the quiet ones that quietly dominate token volume are not.
Pricing the deferrable half
Claude Haiku 4.5 lists at $1.00 per million input tokens and $5.00 per million output. The nightly classification job consumes 40,000 × 1,200 = 48,000,000 input tokens a day, which at $1.00 per million is $48.00, and 40,000 × 150 = 6,000,000 output tokens, which at $5.00 per million is $30.00. That is $78.00 a day, or $2,340.00 over thirty days. Batched at −50%, it is $1,170.00. Saving: $1,170.00 a month.
The weekly summaries run on Claude Sonnet 5 at $2.00/$10.00. Input is 900 × 25,000 = 22,500,000 tokens, or $45.00 at $2.00 per million. Output is 900 × 1,200 = 1,080,000 tokens, or $10.80 at $10.00 per million. That is $55.80 a week — $241.61 a month at 4.33 weeks, and $120.81 batched. Saving: $120.80 a month.
The backfill is a one-off but a large one. Input is 2,400,000 × 1,200 = 2,880,000,000 tokens, which at Haiku 4.5's $1.00 per million is $2,880.00. Output is 2,400,000 × 150 = 360,000,000 tokens, or $1,800.00 at $5.00 per million. Total $4,680.00 at list, $2,340.00 through the batch endpoint. Saving: $2,340.00 each time the taxonomy changes.
Set that against the interactive side, which is untouchable. The chat assistant burns 880,000,000 input tokens ($1,760.00) and 77,000,000 output tokens ($770.00) on Sonnet 5, so $2,530.00. Inline suggestions on Gemini 2.5 Flash-Lite at $0.10/$0.40 come to $186.00 of input and $49.60 of output, so $235.60. The summarise button is $1,080.00 in and $360.00 out on Sonnet 5, so $1,440.00. Interactive total: $4,205.60 a month.
| Feature | Verdict | Monthly, list | Monthly, batched |
|---|---|---|---|
| Chat assistant (Sonnet 5) | Interactive | $2,530.00 | $2,530.00 |
| Inline suggestions (Flash-Lite) | Interactive | $235.60 | $235.60 |
| Summarise button (Sonnet 5) | Interactive | $1,440.00 | $1,440.00 |
| Nightly classification (Haiku 4.5) | Deferrable | $2,340.00 | $1,170.00 |
| Weekly summaries (Sonnet 5) | Deferrable | $241.61 | $120.81 |
| Recurring total | $6,787.21 | $5,496.41 |
A 19.0% cut to the whole bill, from an audit and a change of endpoint. And the summarise button — the feature that could be moved by a product decision rather than an engineering one — is $1,440 a month sitting in the wrong column. Pre-generating summaries for changed threads would drop it into the batch tier and make the button instant. Run your own numbers through the batch API savings calculator before you argue about it.
What you are actually promised
The published commitment across the major vendors is a completion window, not a completion time. OpenAI and Anthropic both target 24 hours; in practice most jobs land in minutes to a few hours, and the small ones often come back faster than the queue depth would suggest. But the number you design against is 24 hours, not the median you observed on a Tuesday. If your nightly job must be on a dashboard by 08:00, submit it at 02:00 and build a fallback that reruns anything still pending at 07:00 through the synchronous endpoint. Do not build a process that breaks when the window is actually used.
Failures are per item, not per batch. A batch is a container, not a transaction. If 40,000 requests go in and 63 hit a content filter, exceed a token limit, or trip a transient error, you get 39,937 results and 63 error records in the output, and you are billed for the ones that completed. There is no rollback and no partial refund, which is the correct behaviour but means your consumer has to handle a results file with holes in it. Write the retry loop on day one: parse the errors, requeue the recoverable ones, and alert on the rest. Teams that skip this discover it during the backfill, at scale, at two in the morning.
Expiry is the other edge. Anything unfinished when the window closes comes back unprocessed and is not billed. That is a refund on the money and a bill on your schedule, so treat window expiry as a real operational event rather than a theoretical one.
Which vendors publish a batch tier
In our 15 August 2026 snapshot, exactly two vendors publish a −50% batch line across their range: OpenAI, on GPT-5.6 Sol, Terra and Luna alike, and Anthropic, on Claude Opus 5, Sonnet 5 and Haiku 4.5. Google, DeepSeek, Mistral, xAI, Cohere, and the open-weight hosts Together and Fireworks do not carry a published batch discount in the index.
That absence changes the shape of the decision more than it looks. The nightly classification job above saves $1,170 a month partly because it happens to run on Anthropic. The same job on Gemini 2.5 Flash at $0.30/$2.50 would cost 48M × $0.30 = $14.40 plus 6M × $2.50 = $15.00, so $29.40 a day and $882.00 a month — cheaper at list than Haiku 4.5 is after batching. The batch discount is a lever on top of a tier choice, not a substitute for one; if you have not yet settled the tier, read choosing a model tier first and apply the discount afterwards.
Where batch earns its unusual status is risk. Switching tiers changes the weights behind your prompt and obliges you to re-run your evals. Trimming prompts changes what the model sees. Caching changes nothing about quality but does change your prompt architecture, since the stable prefix has to come first. Batch changes the transport and nothing else: identical model, identical prompt, identical sampling, identical output distribution. Rolling back is a one-line change to the endpoint you post to. Nothing else on the cost menu is that cheap to try or that cheap to undo — which is why the audit above should be a half-day of work, not a quarter's project.
All figures in this guide come from the price index, last verified 15 August 2026. Every entry links to the vendor page it was read from — see the price index and the methodology.