AI API cost calculator
Enter your real workload — requests per day, tokens per request, cache hit rate — and see the monthly bill for 1,809 priced models. The calculator also shows the cheapest models for your exact profile, so you can see what switching would save.
Prices refreshed 2026-09-13 13:40 UTC. Estimates exclude minimum commitments, tiered discounts and regional multipliers.
Your workload
Estimated cost
Based on 300,000 requests/month with 1,200 prompt + 450 completion tokens. Pricing for GPT 5.5 Pro.
Cheapest models for this exact workload
| Model | $/month | vs you |
|---|---|---|
| Gemini Exp 1114 | free | −100% |
| Gemini Exp 1206 | free | −100% |
| Gemma 3 27b It | free | −100% |
| Gemma 4 26b A4b It | free | −100% |
| Gemma 4 31b It | free | −100% |
| Learnlm 1.5 Pro Experimental | free | −100% |
Prefer a shortlist?
Frequently asked
› How is the monthly cost calculated?
Cost = requests per day × 30 × ((prompt tokens × input price) + (completion tokens × output price)) ÷ 1,000,000, with cached prompt tokens billed at the cached-input rate when a model publishes one.
› Why does my invoice differ from this estimate?
Real bills include minimum commitments, committed-use discounts, regional pricing, retries, failed generations and tool-call overhead. Treat this as a like-for-like comparison between models, not a forecast of an invoice.
› What is a prompt-cache hit rate?
It is the share of prompt tokens that hit the provider prompt cache. OpenAI, Anthropic and Google bill cached tokens at a large discount, so a realistic hit rate can cut the input side of a bill by 50-90%.