AI price index
273 individually sourced prices covering LLM APIs, embedding models, image and video generation, speech, and GPU rental. Each one is re-read from the vendor's own pricing page every week; when a figure moves, the movement is recorded in the change log.
Every model name links to its own page with weekly price history and worked examples — or browse by model and by vendor.
LLM APIs
Price your workload →Per-million-token list prices. Cached input applies to a repeated prompt prefix; batch applies to asynchronous processing.
Embedding models
Price your workload →Embedding prices are small enough that retrieval quality, not price, should drive the choice.
| Model | $/Mtok | Source |
|---|---|---|
| Perplexity pplx-embed-v1-0.6bnew | $0.004 | source |
| Perplexity pplx-embed-context-v1-0.6bnew | $0.008 | source |
| OpenAI text-embedding-3-small | $0.020 | source |
| Voyage AI voyage-4-lite | $0.020 | source |
| Voyage AI rerank-3-litenew | $0.020 | source |
| Voyage AI rerank-2.5-litenew | $0.020 | source |
| Perplexity pplx-embed-v1-4bnew | $0.030 | source |
| Perplexity pplx-embed-context-v1-4bnew | $0.050 | source |
| Voyage AI rerank-3new | $0.050 | source |
| Voyage AI rerank-2.5new | $0.050 | source |
| Voyage AI voyage-4 | $0.060 | source |
| Mistral AI Mistral Embednew | $0.100 | source |
| Voyage AI voyage-4-large | $0.120 | source |
| Voyage AI voyage-code-4 | $0.120 | source |
| Voyage AI voyage-context-4 | $0.120 | source |
| Voyage AI voyage-finance-2new | $0.120 | source |
| Voyage AI voyage-law-2new | $0.120 | source |
| Voyage AI voyage-code-2new | $0.120 | source |
| Voyage AI voyage-multimodal-3.5new | $0.120 | source |
| Voyage AI voyage-multimodal-3new | $0.120 | source |
| Voyage AI voyage-multilingual-2new | $0.120 | source |
| OpenAI text-embedding-3-large | $0.130 | source |
| Mistral AI Codestral Embednew | $0.150 | source |
Image generation
Price your workload →Per-image list prices. Effective cost is this figure multiplied by how many attempts a usable asset takes.
Video generation
Price your workload →Per-second list prices, with a 30-second clip shown because that is where the numbers stop feeling small.
Speech to text
Price your workload →Billed on audio duration including silence, so trimming dead air at ingest is a direct saving.
Text to speech
Price your workload →Note the two different billing units — per character and per minute are not directly comparable without converting.
| Model | List price | Unit | Source |
|---|---|---|---|
| Google TTS Standard/WaveNet | $4.00 | 1M chars | source |
| Deepgram Aura-1 | $15.00 | 1M chars | source |
| xAI Grok Voice APInew | $15.00 | 1M chars | source |
| xAI grok-voicenew | $15.00 | 1M chars | source |
| Google TTS Neural2 | $16.00 | 1M chars | source |
| Mistral AI Voxtral TTSnew | $16.00 | 1M chars | source |
| Google Google TTS Polyglotnew | $16.00 | 1M chars | source |
| Google TTS Chirp 3 HD | $30.00 | 1M chars | source |
| Deepgram Aura-2 | $30.00 | 1M chars | source |
| Deepgram Flux TTSnew | $45.00 | 1M chars | source |
| RunPod Minimax Speech 02 HDnew | $50.00 | 1M chars | source |
| Google Google TTS Instant custom voicenew | $60.00 | 1M chars | source |
| Cartesia Sonic-3new | $65.00 | 1M chars | source |
| Cartesia Sonic-2new | $65.00 | 1M chars | source |
| Google Google TTS Studio voicesnew | $160.00 | 1M chars | source |
| Runway Text to Speechnew | $480.00 | 1M chars | source |
GPU rental
Price your workload →On-demand hourly rates. The monthly column assumes 730 hours — a rented GPU costs that whether you use it or not.
| Vendor | GPU | VRAM | $/hour | $/month at 100% | Type | Source |
|---|---|---|---|---|---|---|
| Vast.ai | L40S | 48 GB | $0.40 | $292 | marketplace | source |
| Vast.ai | A100 SXM4 80GB | 80 GB | $0.52 | $380 | marketplace | source |
| AWS | AWS Trainiumnew | 32 GB | $0.60 | $435 | reserved | source |
| Lambda | Quadro RTX 6000new | 24 GB | $0.69 | $504 | on-demand | source |
| Google Cloud | L4 (G2, per GPU) | 24 GB | $0.71 | $518 | on-demand | source |
| Lambda | V100new | 16 GB | $0.79 | $577 | on-demand | source |
| RunPod | L40S | 48 GB | $1.09 | $796 | on-demand | source |
| Lambda | A6000new | 48 GB | $1.09 | $796 | on-demand | source |
| CoreWeave | L40new | 48 GB | $1.25 | $912 | on-demand | source |
| Lambda | A10new | 24 GB | $1.29 | $942 | on-demand | source |
| AWS | A100new | 80 GB | $1.48 | $1,077 | reserved | source |
| RunPod | A100 80GB PCIe | 80 GB | $1.59 | $1,161 | on-demand | source |
| Vast.ai | H100 SXM | 80 GB | $1.60 | $1,168 | marketplace | source |
| Lambda | A100 SXMnew | 40 GB | $1.99 | $1,453 | on-demand | source |
| Lambda | A100 PCIenew | 40 GB | $1.99 | $1,453 | on-demand | source |
| AWS | AWS Trainium2new | 192 GB | $2.23 | $1,632 | reserved | source |
| Lambda | GH200new | 96 GB | $2.29 | $1,672 | on-demand | source |
| CoreWeave | RTX PRO 6000 Blackwellnew | 96 GB | $2.50 | $1,825 | on-demand | source |
| Lambda Labs | A100 80GB SXM | 80 GB | $2.79 | $2,037 | on-demand | source |
| RunPod | H100 PCIe | 80 GB | $2.89 | $2,110 | on-demand | source |
| fal | RTX PRO 6000new | 96 GB | $2.99 | $2,183 | on-demand | source |
| RunPod | H100 NVLnew | 94 GB | $3.19 | $2,329 | on-demand | source |
| RunPod | H100 SXM | 80 GB | $3.49 | $2,548 | on-demand | source |
| Vast.ai | H200 | 141 GB | $3.82 | $2,789 | marketplace | source |
| Lambda Labs | H100 SXM | 80 GB | $3.99 | $2,913 | on-demand | source |
| Together AI | H100 | 80 GB | $3.99 | $2,913 | on-demand | source |
| Together AI | HGX H100new | 80 GB | $3.99 | $2,913 | on-demand | source |
| Vast.ai | B200 | 192 GB | $4.13 | $3,015 | marketplace | source |
| CoreWeave | HGX B300new | 270 GB | $4.48 | $3,270 | spot | source |
| RunPod | H200 | 141 GB | $4.59 | $3,351 | on-demand | source |
| AWS | H100 (p5.48xlarge, per GPU) | 80 GB | $5.19 | $3,789 | capacity block | source |
| Together AI | HGX H200new | 141 GB | $5.99 | $4,373 | on-demand | source |
| CoreWeave | H100 HGX (per GPU) | 80 GB | $6.16 | $4,497 | on-demand | source |
| CoreWeave | H200 HGXnew | 141 GB | $6.31 | $4,606 | on-demand | source |
| Lambda Labs | B200 SXM6 | 180 GB | $6.69 | $4,884 | on-demand | source |
| RunPod | B200 | 180 GB | $6.79 | $4,957 | on-demand | source |
| Google Cloud | B200 (A4 High, per GPU) | 180 GB | $8.06 | $5,884 | on-demand | source |
| Together AI | B200 | 180 GB | $8.19 | $5,979 | on-demand | source |
| CoreWeave | B200 HGX (per GPU) | 180 GB | $8.60 | $6,278 | on-demand | source |
| CoreWeave | GB200 NVL72new | 186 GB | $10.50 | $7,665 | on-demand | source |
| AWS | B200 (p6-b200, per GPU) | 180 GB | $12.36 | $9,019 | capacity block | source |
| AWS | B300new | 192 GB | $14.04 | $10,249 | reserved | source |
Using these figures
Every number here is a published list price in US dollars, excluding tax. If your organisation has a volume commitment or a negotiated agreement, your real rate is lower and none of these figures apply directly.
To turn a price into a monthly bill you need a workload, which is what the calculators are for. To understand why the input and output columns differ by a factor of five or more, see why output tokens cost more.
The methodology page documents how this index is built, how failures to verify are handled, and which vendors publish prices in a form that cannot be read reliably.