The Magic of AI

AI price index

101 individually sourced prices covering LLM APIs, embedding models, image and video generation, speech, and GPU rental. Each one is re-read from the vendor's own pricing page every week; when a figure moves, the movement is recorded in the change log.

Every figure links to its source. Prices are USD list rates excluding tax, read from each vendor's own pricing page. Negotiated rates, volume commitments and credits are not modelled. Click a column header to sort.

LLM APIs

Per-million-token list prices. Cached input applies to a repeated prompt prefix; batch applies to asynchronous processing.

ModelTierInput $/MtokOutput $/MtokCached inputBatchContextSource
Mistral AI Ministral 3 (3B)budget$0.10$0.10source
Google Gemini 2.5 Flash-Litebudget$0.10$0.40$0.010−50%source
DeepSeek DeepSeek V4 Flashbudget$0.14$0.28$0.0031000ksource
Meta via Fireworks AI Llama 4 Scout Instructsmall$0.15$0.601048ksource
Mistral AI Mistral Small 4small$0.15$0.60source
Qwen via Together AI Qwen3 235B A22B Instructsmall$0.20$0.60source
Qwen via Together AI Qwen2.5 7B Instruct Turbobudget$0.30$0.30source
Cohere Command-lightbudget$0.30$0.60source
OpenAI GPT-5.6 Lunasmall$0.20$1.20$0.020−50%1050ksource
Mistral AI Codestralsmall$0.30$0.90source
DeepSeek DeepSeek V4 Prosmall$0.43$0.87$0.0041000ksource
Mistral AI Mistral Large 3mid$0.50$1.50source
Cohere Command R (03-2024)small$0.50$1.50source
Google Gemini 2.5 Flashsmall$0.30$2.50$0.030−50%source
Google Gemini 3.5 Flash-Litesmall$0.30$2.50$0.030source
Meta via Together AI Llama 3.3 70Bmid$1.04$1.04source
Anthropic Claude Haiku 4.5small$1.00$5.00$0.100−50%200ksource
Mistral AI Mistral Medium 3.5mid$1.50$7.50source
xAI Grok 4.6mid$2.00$6.00500ksource
Google Gemini 2.5 Promid$1.25$10.00$0.125−50%source
Google Gemini 3.5 Flashmid$1.50$9.00$0.150−50%source
Anthropic Claude Sonnet 5mid$2.00$10.00$0.200−50%1000ksource
OpenAI GPT-5.6 Terramid$2.00$12.00$0.200−50%1050ksource
Google Gemini 3.1 Profrontier$2.00$12.00−50%source
Cohere Command R+ (08-2024)mid$2.50$10.00source
Anthropic Claude Opus 5frontier$5.00$25.00$0.500−50%1000ksource
OpenAI GPT-5.6 Solfrontier$5.00$30.00$0.500−50%1050ksource

Embedding models

Embedding prices are small enough that retrieval quality, not price, should drive the choice.

Model$/MtokSource
OpenAI text-embedding-3-small$0.020source
Voyage AI voyage-4-lite$0.020source
Voyage AI voyage-4$0.060source
Voyage AI voyage-4-large$0.120source
Voyage AI voyage-code-4$0.120source
Voyage AI voyage-context-4$0.120source
OpenAI text-embedding-3-large$0.130source

Image generation

Per-image list prices. Effective cost is this figure multiplied by how many attempts a usable asset takes.

ModelTier$/imageSource
Black Forest Labs FLUX.2 Klein 4B1MP$0.014source
Kling AI Kling Image 2.11K/2K$0.014source
Google Imagen 4 Fastfast$0.020source
Stability AI Stable Diffusion 3.5 Flashflash$0.025source
Kling AI Kling Image 3.01K/2K$0.028source
Black Forest Labs FLUX.2 Pro1MP$0.030source
Stability AI Stable Image Corecost-optimised$0.030source
Google Imagen 4standard$0.040source
Black Forest Labs FLUX 1.1 [pro]standard$0.040source
Google Imagen 4 Ultraultra$0.060source
Stability AI Stable Diffusion 3.5 Largestandard$0.065source
Black Forest Labs FLUX.2 Max1MP$0.070source
Black Forest Labs FLUX Kontext [max]max$0.080source
Stability AI Stable Image Ultraflagship$0.080source

Video generation

Per-second list prices, with a 30-second clip shown because that is where the numbers stop feeling small.

ModelTier$/second$/30s clipSource
Kling AI Kling 2.5 Turbo720p$0.042$1.26source
Google Veo 3.1 Lite720p + audio$0.050$1.50source
Luma AI Ray3.2720p$0.060$1.80source
Kling AI Kling 3.0720p$0.084$2.52source
Google Veo 3.1 Fast720p + audio$0.100$3.00source
Kling AI Kling 3.0 Turbo720p + audio$0.112$3.36source
Black Forest Labs FLUX 3 VideoHD full render$0.170$5.10source
Luma AI Ray3.21080p$0.240$7.20source
Google Veo 3.1720p/1080p + audio$0.400$12.00source
Google Veo 3.14K + audio$0.600$18.00source

Speech to text

Billed on audio duration including silence, so trimming dead air at ingest is a direct saving.

Model$/minute$/audio hourSource
AssemblyAI Universal-2$0.0025$0.15source
OpenAI gpt-4o-mini-transcribe$0.0030$0.18source
Google Speech-to-Text V2 Dynamic Batch$0.0030$0.18source
AssemblyAI Universal-3.5 Pro$0.0035$0.21source
OpenAI gpt-transcribe$0.0045$0.27source
Deepgram Nova-3 monolingual$0.0048$0.29source
Deepgram Nova-3 multilingual$0.0058$0.35source
OpenAI gpt-4o-transcribe$0.0060$0.36source
Google Speech-to-Text V2 Standard$0.0160$0.96source
ElevenLabs Scribe v2 (Pro rate)$0.0545$3.27source

Text to speech

Note the two different billing units — per character and per minute are not directly comparable without converting.

ModelList priceUnitSource
Google TTS Standard/WaveNet$4.001M charssource
Deepgram Aura-1$15.001M charssource
Google TTS Neural2$16.001M charssource
Google TTS Chirp 3 HD$30.001M charssource
Deepgram Aura-2$30.001M charssource

GPU rental

On-demand hourly rates. The monthly column assumes 730 hours — a rented GPU costs that whether you use it or not.

VendorGPUVRAM$/hour$/month at 100%TypeSource
Vast.aiL40S48 GB$0.40$292marketplacesource
Vast.aiA100 SXM4 80GB80 GB$0.52$380marketplacesource
Google CloudL4 (G2, per GPU)24 GB$0.71$518on-demandsource
RunPodL40S48 GB$0.99$723on-demandsource
RunPodA100 80GB PCIe80 GB$1.39$1,015on-demandsource
Vast.aiH100 SXM80 GB$1.60$1,168marketplacesource
Lambda LabsA100 80GB SXM80 GB$2.79$2,037on-demandsource
RunPodH100 PCIe80 GB$2.89$2,110on-demandsource
RunPodH100 SXM80 GB$3.29$2,402on-demandsource
Vast.aiH200141 GB$3.82$2,789marketplacesource
Lambda LabsH100 SXM80 GB$3.99$2,913on-demandsource
Together AIH10080 GB$3.99$2,913on-demandsource
Vast.aiB200192 GB$4.13$3,015marketplacesource
RunPodH200141 GB$4.59$3,351on-demandsource
AWSH100 (p5.48xlarge, per GPU)80 GB$5.19$3,789capacity blocksource
CoreWeaveH100 HGX (per GPU)80 GB$6.16$4,493on-demandsource
Lambda LabsB200 SXM6180 GB$6.69$4,884on-demandsource
RunPodB200180 GB$6.79$4,957on-demandsource
Google CloudB200 (A4 High, per GPU)180 GB$8.06$5,884on-demandsource
Together AIB200180 GB$8.19$5,979on-demandsource
CoreWeaveB200 HGX (per GPU)180 GB$8.60$6,278on-demandsource
AWSB200 (p6-b200, per GPU)180 GB$12.36$9,019capacity blocksource

Using these figures

Every number here is a published list price in US dollars, excluding tax. If your organisation has a volume commitment or a negotiated agreement, your real rate is lower and none of these figures apply directly.

To turn a price into a monthly bill you need a workload, which is what the calculators are for. To understand why the input and output columns differ by a factor of five or more, see why output tokens cost more.

The methodology page documents how this index is built, how failures to verify are handled, and which vendors publish prices in a form that cannot be read reliably.