Price per 1M logo

Cheapest cached-input pricing

Models with explicit cache-read pricing — the single biggest lever for high-volume apps that resend the same system prompt.

Current leader: Nemotron 3.5 Lightning 30b A3b (Perplexity). Refreshed 2026-09-13 13:40 UTC.

#ModelProviderInput $/1MOutput $/1MCached $/1MContext
1Nemotron 3.5 Lightning 30b A3bPerplexity$0.011$0.17$0.0011
2Muse Spark 1.2 ContributorMeta$0.1$0.2$0.0021.0M
3Muse Spark 1.3 ContributorMeta$0.1$0.2$0.0021.0M
4Mimo V2.5OpenRouter$0.14$0.28$0.00281.0M
5Deepseek V4.1 FlashOpenRouter$0.15$0.6$0.0031.0M
6Mimo V2.5Novita AI$0.168$0.336$0.00341.0M
7Mimo V2.5 ProOpenRouter$0.435$0.87$0.00361.0M
8Mimo V2.5 ProNovita AI$0.522$1.04$0.00431.0M
9Databricks GPT 5 NanoDatabricks$0.05$0.4$0.005272K
10GPT 5 NanoOpenAI$0.05$0.4$0.005272K
11GPT 5 NanoOpenAI$0.05$0.4$0.005272K
12GPT 5 NanoAzure OpenAI$0.05$0.4$0.005272K
13GPT 5 NanoOpenRouter$0.05$0.4$0.005272K
14GPT 5 NanoAzure OpenAI$0.055$0.44$0.0055272K
15Deepseek FlashDeepSeek$0.3$1.2$0.0061M
16Deepseek V4 FlashDeepSeek$0.3$1.2$0.0061M
17Deepseek V4 Flash Vision ExpDeepSeek$0.3$1.2$0.0061M
18Qwen3.7 FlashOpenRouter$0.03$0.13$0.0061M
19Deepseek V4 Flash 0731Fireworks AI$0.22$0.66$0.0071.0M
20Deepseek V4p1 FlashFireworks AI$0.22$0.66$0.0071.0M
21Deepseek V4 Flash Vision ExpFireworks AI$0.22$0.66$0.0071.0M
22Deepseek V4 Flash Vision ExpOpenRouter$0.22$0.66$0.0071.0M
23Laguna S 2.1OpenRouter$0.09$0.18$0.0091.0M
24Gemini 2.5 Flash LiteGoogle Gemini$0.1$0.4$0.011.0M
25Gemini 2.5 Flash Lite Preview 09 2025Google Gemini$0.1$0.4$0.011.0M
26Gemini Flash LiteGoogle Gemini$0.1$0.4$0.011.0M
27Gemini 2.5 Flash Lite Preview 06 17Google Gemini$0.1$0.4$0.011.0M
28Ministral 3b 2512Mistral AI$0.1$0.1$0.01131K
29Ministral 3bMistral AI$0.1$0.1$0.01131K
30Ministral 3 3b 2512Mistral AI$0.1$0.1$0.01131K
31GLM 4.7 FlashDeepInfra$0.06$0.4$0.01203K
32Nemotron Lightning 3p5 30b A3bFireworks AI$0.05$0.2$0.01262K
33Glm 4.7 FlashNovita AI$0.07$0.4$0.01200K
34Mimo V2 FlashOpenRouter$0.1$0.3$0.01262K
35Gemini 2.5 Flash LiteOpenRouter$0.1$0.4$0.011.0M
36Voxtral Small 24b 2507OpenRouter$0.1$0.3$0.0133K
37Glm 4.7 FlashOpenRouter$0.06$0.4$0.01200K
38Ling 3.0 FlashDeepInfra$0.06$0.18$0.012131K
39Ling 3.0 Flash FastNovita AI$0.06$0.18$0.012262K
40Ling 3.0 FlashNovita AI$0.06$0.18$0.012262K
41Magistral SmallMistral AI$0.15$0.6$0.015262K
42Mistral SmallMistral AI$0.15$0.6$0.015262K
43Ministral 3 8b 2512Mistral AI$0.15$0.15$0.015262K
44Ministral 8b 2512Mistral AI$0.15$0.15$0.015262K
45Ministral 8bMistral AI$0.15$0.15$0.015262K
46Mistral Small 2603Mistral AI$0.15$0.6$0.015262K
47Mistral Vibe Cli FastMistral AI$0.15$0.6$0.015262K
48GPT Oss 120bFireworks AI$0.15$0.6$0.015131K
49Mistral Small 2603OpenRouter$0.15$0.6$0.015262K
50DeepSeek V4 Flash 0731DeepInfra$0.08$0.18$0.0161.0M
51Qwen3.8 FlashOpenRouter$0.15$0.47$0.0161M
52Deepseek V4 Flash 0731OpenRouter$0.065$0.18$0.0161.3M
53Deepseek V4 FlashOpenRouter$0.085$0.171$0.0171.0M
54DeepSeek V4 FlashDeepInfra$0.09$0.18$0.0181.0M
55Gemini 2.0 Flash LiteGoogle Gemini$0.075$0.3$0.0191.0M
56Gemini 2.0 Flash Lite 001Google Gemini$0.075$0.3$0.0191.0M
57Deepseek V4 Pro 0813OpenRouter$0.579$1.74$0.0191.0M
58Ministral 14b 2512Mistral AI$0.2$0.2$0.02262K
59Ministral 14bMistral AI$0.2$0.2$0.02262K
60Ministral 3 14b 2512Mistral AI$0.2$0.2$0.02262K

Other use cases

Frequently asked

Which provider has the cheapest prompt caching?

Nemotron 3.5 Lightning 30b A3b from Perplexity currently leads this list, at $0.0115/1M input tokens. Rankings are recomputed from the raw dataset on every refresh, so they track price cuts automatically.

How are these rankings calculated?

Each use case applies a capability filter (context size, vision, tool calling, cache pricing) and then ranks the surviving models by the blended cost that profile actually pays. See the methodology page for the exact formulas.

Is the cheapest model always the right choice?

No. Check the capability flags and context window, and benchmark quality on your own data — the cheapest model that fails your evals costs more than the second cheapest that passes them.