Cheapest cached-input pricing
Models with explicit cache-read pricing — the single biggest lever for high-volume apps that resend the same system prompt.
Current leader: Nemotron 3.5 Lightning 30b A3b (Perplexity). Refreshed 2026-09-13 13:40 UTC.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Nemotron 3.5 Lightning 30b A3b | Perplexity | $0.011 | $0.17 | $0.0011 | — |
| 2 | Muse Spark 1.2 Contributor | Meta | $0.1 | $0.2 | $0.002 | 1.0M |
| 3 | Muse Spark 1.3 Contributor | Meta | $0.1 | $0.2 | $0.002 | 1.0M |
| 4 | Mimo V2.5 | OpenRouter | $0.14 | $0.28 | $0.0028 | 1.0M |
| 5 | Deepseek V4.1 Flash | OpenRouter | $0.15 | $0.6 | $0.003 | 1.0M |
| 6 | Mimo V2.5 | Novita AI | $0.168 | $0.336 | $0.0034 | 1.0M |
| 7 | Mimo V2.5 Pro | OpenRouter | $0.435 | $0.87 | $0.0036 | 1.0M |
| 8 | Mimo V2.5 Pro | Novita AI | $0.522 | $1.04 | $0.0043 | 1.0M |
| 9 | Databricks GPT 5 Nano | Databricks | $0.05 | $0.4 | $0.005 | 272K |
| 10 | GPT 5 Nano | OpenAI | $0.05 | $0.4 | $0.005 | 272K |
| 11 | GPT 5 Nano | OpenAI | $0.05 | $0.4 | $0.005 | 272K |
| 12 | GPT 5 Nano | Azure OpenAI | $0.05 | $0.4 | $0.005 | 272K |
| 13 | GPT 5 Nano | OpenRouter | $0.05 | $0.4 | $0.005 | 272K |
| 14 | GPT 5 Nano | Azure OpenAI | $0.055 | $0.44 | $0.0055 | 272K |
| 15 | Deepseek Flash | DeepSeek | $0.3 | $1.2 | $0.006 | 1M |
| 16 | Deepseek V4 Flash | DeepSeek | $0.3 | $1.2 | $0.006 | 1M |
| 17 | Deepseek V4 Flash Vision Exp | DeepSeek | $0.3 | $1.2 | $0.006 | 1M |
| 18 | Qwen3.7 Flash | OpenRouter | $0.03 | $0.13 | $0.006 | 1M |
| 19 | Deepseek V4 Flash 0731 | Fireworks AI | $0.22 | $0.66 | $0.007 | 1.0M |
| 20 | Deepseek V4p1 Flash | Fireworks AI | $0.22 | $0.66 | $0.007 | 1.0M |
| 21 | Deepseek V4 Flash Vision Exp | Fireworks AI | $0.22 | $0.66 | $0.007 | 1.0M |
| 22 | Deepseek V4 Flash Vision Exp | OpenRouter | $0.22 | $0.66 | $0.007 | 1.0M |
| 23 | Laguna S 2.1 | OpenRouter | $0.09 | $0.18 | $0.009 | 1.0M |
| 24 | Gemini 2.5 Flash Lite | Google Gemini | $0.1 | $0.4 | $0.01 | 1.0M |
| 25 | Gemini 2.5 Flash Lite Preview 09 2025 | Google Gemini | $0.1 | $0.4 | $0.01 | 1.0M |
| 26 | Gemini Flash Lite | Google Gemini | $0.1 | $0.4 | $0.01 | 1.0M |
| 27 | Gemini 2.5 Flash Lite Preview 06 17 | Google Gemini | $0.1 | $0.4 | $0.01 | 1.0M |
| 28 | Ministral 3b 2512 | Mistral AI | $0.1 | $0.1 | $0.01 | 131K |
| 29 | Ministral 3b | Mistral AI | $0.1 | $0.1 | $0.01 | 131K |
| 30 | Ministral 3 3b 2512 | Mistral AI | $0.1 | $0.1 | $0.01 | 131K |
| 31 | GLM 4.7 Flash | DeepInfra | $0.06 | $0.4 | $0.01 | 203K |
| 32 | Nemotron Lightning 3p5 30b A3b | Fireworks AI | $0.05 | $0.2 | $0.01 | 262K |
| 33 | Glm 4.7 Flash | Novita AI | $0.07 | $0.4 | $0.01 | 200K |
| 34 | Mimo V2 Flash | OpenRouter | $0.1 | $0.3 | $0.01 | 262K |
| 35 | Gemini 2.5 Flash Lite | OpenRouter | $0.1 | $0.4 | $0.01 | 1.0M |
| 36 | Voxtral Small 24b 2507 | OpenRouter | $0.1 | $0.3 | $0.01 | 33K |
| 37 | Glm 4.7 Flash | OpenRouter | $0.06 | $0.4 | $0.01 | 200K |
| 38 | Ling 3.0 Flash | DeepInfra | $0.06 | $0.18 | $0.012 | 131K |
| 39 | Ling 3.0 Flash Fast | Novita AI | $0.06 | $0.18 | $0.012 | 262K |
| 40 | Ling 3.0 Flash | Novita AI | $0.06 | $0.18 | $0.012 | 262K |
| 41 | Magistral Small | Mistral AI | $0.15 | $0.6 | $0.015 | 262K |
| 42 | Mistral Small | Mistral AI | $0.15 | $0.6 | $0.015 | 262K |
| 43 | Ministral 3 8b 2512 | Mistral AI | $0.15 | $0.15 | $0.015 | 262K |
| 44 | Ministral 8b 2512 | Mistral AI | $0.15 | $0.15 | $0.015 | 262K |
| 45 | Ministral 8b | Mistral AI | $0.15 | $0.15 | $0.015 | 262K |
| 46 | Mistral Small 2603 | Mistral AI | $0.15 | $0.6 | $0.015 | 262K |
| 47 | Mistral Vibe Cli Fast | Mistral AI | $0.15 | $0.6 | $0.015 | 262K |
| 48 | GPT Oss 120b | Fireworks AI | $0.15 | $0.6 | $0.015 | 131K |
| 49 | Mistral Small 2603 | OpenRouter | $0.15 | $0.6 | $0.015 | 262K |
| 50 | DeepSeek V4 Flash 0731 | DeepInfra | $0.08 | $0.18 | $0.016 | 1.0M |
| 51 | Qwen3.8 Flash | OpenRouter | $0.15 | $0.47 | $0.016 | 1M |
| 52 | Deepseek V4 Flash 0731 | OpenRouter | $0.065 | $0.18 | $0.016 | 1.3M |
| 53 | Deepseek V4 Flash | OpenRouter | $0.085 | $0.171 | $0.017 | 1.0M |
| 54 | DeepSeek V4 Flash | DeepInfra | $0.09 | $0.18 | $0.018 | 1.0M |
| 55 | Gemini 2.0 Flash Lite | Google Gemini | $0.075 | $0.3 | $0.019 | 1.0M |
| 56 | Gemini 2.0 Flash Lite 001 | Google Gemini | $0.075 | $0.3 | $0.019 | 1.0M |
| 57 | Deepseek V4 Pro 0813 | OpenRouter | $0.579 | $1.74 | $0.019 | 1.0M |
| 58 | Ministral 14b 2512 | Mistral AI | $0.2 | $0.2 | $0.02 | 262K |
| 59 | Ministral 14b | Mistral AI | $0.2 | $0.2 | $0.02 | 262K |
| 60 | Ministral 3 14b 2512 | Mistral AI | $0.2 | $0.2 | $0.02 | 262K |
Other use cases
Frequently asked
› Which provider has the cheapest prompt caching?
Nemotron 3.5 Lightning 30b A3b from Perplexity currently leads this list, at $0.0115/1M input tokens. Rankings are recomputed from the raw dataset on every refresh, so they track price cuts automatically.
› How are these rankings calculated?
Each use case applies a capability filter (context size, vision, tool calling, cache pricing) and then ranks the surviving models by the blended cost that profile actually pays. See the methodology page for the exact formulas.
› Is the cheapest model always the right choice?
No. Check the capability flags and context window, and benchmark quality on your own data — the cheapest model that fails your evals costs more than the second cheapest that passes them.