The visible answer says 500 tokens. The invoice says 4,500. The difference is reasoning — the model's internal working, generated before the answer, billed at the output rate, and never returned to you. This calculator makes that gap explicit and prices it across every model at once.
Reasoning tokens are output tokens you cannot read
When a reasoning model "thinks," it generates tokens exactly as it does when it writes an answer. They occupy the same output pathway, they are billed at the same output rate, and the only difference is that the API strips them before responding. So a request billed as input + (visible + reasoning) output tokens can cost five or ten times what the printed answer implies. The "reasoning share of output" stat is the blunt version of the problem: at 4,000 reasoning against 500 visible tokens, 89% of every output dollar is spent on text you never see.
The tax multiplies the output price, so the output price is the lever
Because reasoning is billed as output, its cost scales with each model's output rate — the most volatile number on any price card. On a $30/Mtok output model, 4,000 reasoning tokens add $0.12 per request; on a $0.40/Mtok model they add $0.0016, a 75× difference for the identical thinking budget. The "reasoning tax" column shows how much a request grows once thinking is included, and the table sorts by monthly cost with reasoning on, so you can see which models keep the habit cheap. For reasoning-heavy work, a low output price is worth far more than it is for ordinary generation, because the tax multiplies it on every call.
Sizing a realistic reasoning budget
The default of 4,000 is a middle setting. In practice the budget tracks the effort level you request: a light setting spends roughly 500–1,500 reasoning tokens, a medium one a few thousand, and a high or "extended thinking" setting can run past 8,000 or 16,000 on a hard problem. If your model exposes a reasoning-budget or effort parameter, that cap is the single most effective control you have — most tasks do not need the ceiling, and the marginal reasoning tokens past the point of a correct answer are pure cost. Enter the budget you actually use, not the maximum the model allows.
What this leaves out, and where it connects
This prices one request shape. Two adjustments make it match a real bill. First, reasoning budgets vary by input: hard prompts think longer, so a workload with a mix of easy and hard requests wants a weighted average, not the worst case. Second, if reasoning runs inside an agent loop, the thinking cost lands on every step, compounding with the context growth priced in the AI agent cost calculator. To see how the input side of the same request stacks up across models, the LLM API cost calculator is the companion to this one, and why output tokens cost more explains the asymmetry that makes reasoning expensive in the first place.
Prices used by this calculator were last verified on 31 August 2026 from the vendors' own pricing pages. See the full price index for every figure and its source, the change log for what has moved recently, and the methodology for how the index is maintained and where its limits are.