AI price index
101 individually sourced prices covering LLM APIs, embedding models, image and video generation, speech, and GPU rental. Each one is re-read from the vendor's own pricing page every week; when a figure moves, the movement is recorded in the change log.
LLM APIs
Per-million-token list prices. Cached input applies to a repeated prompt prefix; batch applies to asynchronous processing.
| Model | Tier | Input $/Mtok | Output $/Mtok | Cached input | Batch | Context | Source |
|---|---|---|---|---|---|---|---|
| Mistral AI Ministral 3 (3B) | budget | $0.10 | $0.10 | — | — | — | source |
| Google Gemini 2.5 Flash-Lite | budget | $0.10 | $0.40 | $0.010 | −50% | — | source |
| DeepSeek DeepSeek V4 Flash | budget | $0.14 | $0.28 | $0.003 | — | 1000k | source |
| Meta via Fireworks AI Llama 4 Scout Instruct | small | $0.15 | $0.60 | — | — | 1048k | source |
| Mistral AI Mistral Small 4 | small | $0.15 | $0.60 | — | — | — | source |
| Qwen via Together AI Qwen3 235B A22B Instruct | small | $0.20 | $0.60 | — | — | — | source |
| Qwen via Together AI Qwen2.5 7B Instruct Turbo | budget | $0.30 | $0.30 | — | — | — | source |
| Cohere Command-light | budget | $0.30 | $0.60 | — | — | — | source |
| OpenAI GPT-5.6 Luna | small | $0.20 | $1.20 | $0.020 | −50% | 1050k | source |
| Mistral AI Codestral | small | $0.30 | $0.90 | — | — | — | source |
| DeepSeek DeepSeek V4 Pro | small | $0.43 | $0.87 | $0.004 | — | 1000k | source |
| Mistral AI Mistral Large 3 | mid | $0.50 | $1.50 | — | — | — | source |
| Cohere Command R (03-2024) | small | $0.50 | $1.50 | — | — | — | source |
| Google Gemini 2.5 Flash | small | $0.30 | $2.50 | $0.030 | −50% | — | source |
| Google Gemini 3.5 Flash-Lite | small | $0.30 | $2.50 | $0.030 | — | — | source |
| Meta via Together AI Llama 3.3 70B | mid | $1.04 | $1.04 | — | — | — | source |
| Anthropic Claude Haiku 4.5 | small | $1.00 | $5.00 | $0.100 | −50% | 200k | source |
| Mistral AI Mistral Medium 3.5 | mid | $1.50 | $7.50 | — | — | — | source |
| xAI Grok 4.6 | mid | $2.00 | $6.00 | — | — | 500k | source |
| Google Gemini 2.5 Pro | mid | $1.25 | $10.00 | $0.125 | −50% | — | source |
| Google Gemini 3.5 Flash | mid | $1.50 | $9.00 | $0.150 | −50% | — | source |
| Anthropic Claude Sonnet 5 | mid | $2.00 | $10.00 | $0.200 | −50% | 1000k | source |
| OpenAI GPT-5.6 Terra | mid | $2.00 | $12.00 | $0.200 | −50% | 1050k | source |
| Google Gemini 3.1 Pro | frontier | $2.00 | $12.00 | — | −50% | — | source |
| Cohere Command R+ (08-2024) | mid | $2.50 | $10.00 | — | — | — | source |
| Anthropic Claude Opus 5 | frontier | $5.00 | $25.00 | $0.500 | −50% | 1000k | source |
| OpenAI GPT-5.6 Sol | frontier | $5.00 | $30.00 | $0.500 | −50% | 1050k | source |
Embedding models
Embedding prices are small enough that retrieval quality, not price, should drive the choice.
Image generation
Per-image list prices. Effective cost is this figure multiplied by how many attempts a usable asset takes.
| Model | Tier | $/image | Source |
|---|---|---|---|
| Black Forest Labs FLUX.2 Klein 4B | 1MP | $0.014 | source |
| Kling AI Kling Image 2.1 | 1K/2K | $0.014 | source |
| Google Imagen 4 Fast | fast | $0.020 | source |
| Stability AI Stable Diffusion 3.5 Flash | flash | $0.025 | source |
| Kling AI Kling Image 3.0 | 1K/2K | $0.028 | source |
| Black Forest Labs FLUX.2 Pro | 1MP | $0.030 | source |
| Stability AI Stable Image Core | cost-optimised | $0.030 | source |
| Google Imagen 4 | standard | $0.040 | source |
| Black Forest Labs FLUX 1.1 [pro] | standard | $0.040 | source |
| Google Imagen 4 Ultra | ultra | $0.060 | source |
| Stability AI Stable Diffusion 3.5 Large | standard | $0.065 | source |
| Black Forest Labs FLUX.2 Max | 1MP | $0.070 | source |
| Black Forest Labs FLUX Kontext [max] | max | $0.080 | source |
| Stability AI Stable Image Ultra | flagship | $0.080 | source |
Video generation
Per-second list prices, with a 30-second clip shown because that is where the numbers stop feeling small.
| Model | Tier | $/second | $/30s clip | Source |
|---|---|---|---|---|
| Kling AI Kling 2.5 Turbo | 720p | $0.042 | $1.26 | source |
| Google Veo 3.1 Lite | 720p + audio | $0.050 | $1.50 | source |
| Luma AI Ray3.2 | 720p | $0.060 | $1.80 | source |
| Kling AI Kling 3.0 | 720p | $0.084 | $2.52 | source |
| Google Veo 3.1 Fast | 720p + audio | $0.100 | $3.00 | source |
| Kling AI Kling 3.0 Turbo | 720p + audio | $0.112 | $3.36 | source |
| Black Forest Labs FLUX 3 Video | HD full render | $0.170 | $5.10 | source |
| Luma AI Ray3.2 | 1080p | $0.240 | $7.20 | source |
| Google Veo 3.1 | 720p/1080p + audio | $0.400 | $12.00 | source |
| Google Veo 3.1 | 4K + audio | $0.600 | $18.00 | source |
Speech to text
Billed on audio duration including silence, so trimming dead air at ingest is a direct saving.
| Model | $/minute | $/audio hour | Source |
|---|---|---|---|
| AssemblyAI Universal-2 | $0.0025 | $0.15 | source |
| OpenAI gpt-4o-mini-transcribe | $0.0030 | $0.18 | source |
| Google Speech-to-Text V2 Dynamic Batch | $0.0030 | $0.18 | source |
| AssemblyAI Universal-3.5 Pro | $0.0035 | $0.21 | source |
| OpenAI gpt-transcribe | $0.0045 | $0.27 | source |
| Deepgram Nova-3 monolingual | $0.0048 | $0.29 | source |
| Deepgram Nova-3 multilingual | $0.0058 | $0.35 | source |
| OpenAI gpt-4o-transcribe | $0.0060 | $0.36 | source |
| Google Speech-to-Text V2 Standard | $0.0160 | $0.96 | source |
| ElevenLabs Scribe v2 (Pro rate) | $0.0545 | $3.27 | source |
Text to speech
Note the two different billing units — per character and per minute are not directly comparable without converting.
GPU rental
On-demand hourly rates. The monthly column assumes 730 hours — a rented GPU costs that whether you use it or not.
| Vendor | GPU | VRAM | $/hour | $/month at 100% | Type | Source |
|---|---|---|---|---|---|---|
| Vast.ai | L40S | 48 GB | $0.40 | $292 | marketplace | source |
| Vast.ai | A100 SXM4 80GB | 80 GB | $0.52 | $380 | marketplace | source |
| Google Cloud | L4 (G2, per GPU) | 24 GB | $0.71 | $518 | on-demand | source |
| RunPod | L40S | 48 GB | $0.99 | $723 | on-demand | source |
| RunPod | A100 80GB PCIe | 80 GB | $1.39 | $1,015 | on-demand | source |
| Vast.ai | H100 SXM | 80 GB | $1.60 | $1,168 | marketplace | source |
| Lambda Labs | A100 80GB SXM | 80 GB | $2.79 | $2,037 | on-demand | source |
| RunPod | H100 PCIe | 80 GB | $2.89 | $2,110 | on-demand | source |
| RunPod | H100 SXM | 80 GB | $3.29 | $2,402 | on-demand | source |
| Vast.ai | H200 | 141 GB | $3.82 | $2,789 | marketplace | source |
| Lambda Labs | H100 SXM | 80 GB | $3.99 | $2,913 | on-demand | source |
| Together AI | H100 | 80 GB | $3.99 | $2,913 | on-demand | source |
| Vast.ai | B200 | 192 GB | $4.13 | $3,015 | marketplace | source |
| RunPod | H200 | 141 GB | $4.59 | $3,351 | on-demand | source |
| AWS | H100 (p5.48xlarge, per GPU) | 80 GB | $5.19 | $3,789 | capacity block | source |
| CoreWeave | H100 HGX (per GPU) | 80 GB | $6.16 | $4,493 | on-demand | source |
| Lambda Labs | B200 SXM6 | 180 GB | $6.69 | $4,884 | on-demand | source |
| RunPod | B200 | 180 GB | $6.79 | $4,957 | on-demand | source |
| Google Cloud | B200 (A4 High, per GPU) | 180 GB | $8.06 | $5,884 | on-demand | source |
| Together AI | B200 | 180 GB | $8.19 | $5,979 | on-demand | source |
| CoreWeave | B200 HGX (per GPU) | 180 GB | $8.60 | $6,278 | on-demand | source |
| AWS | B200 (p6-b200, per GPU) | 180 GB | $12.36 | $9,019 | capacity block | source |
Using these figures
Every number here is a published list price in US dollars, excluding tax. If your organisation has a volume commitment or a negotiated agreement, your real rate is lower and none of these figures apply directly.
To turn a price into a monthly bill you need a workload, which is what the calculators are for. To understand why the input and output columns differ by a factor of five or more, see why output tokens cost more.
The methodology page documents how this index is built, how failures to verify are handled, and which vendors publish prices in a form that cannot be read reliably.