Price per 1M logo

Cheapest long-context models

Models with at least a 200K-token context window, ranked by input price — for summarising books, legal dumps and codebases.

Current leader: Gemini Exp 1114 (Google Gemini). Refreshed 2026-09-13 13:40 UTC.

#ModelProviderInput $/1MOutput $/1MCached $/1MContext
1Gemini Exp 1114Google Geminifreefree1.0M
2Gemini Exp 1206Google Geminifreefree2.1M
3Gemma 4 26b A4b ItGoogle Geminifreefree262K
4Gemma 4 31b ItGoogle Geminifreefree262K
5Labs Leanstral 1 5Mistral AIfreefree262K
6Labs Leanstral 1 5 1Mistral AIfreefree262K
7Anthropic.claude Mythos PreviewAmazon Bedrockfreefree1M
8Qwen3 Coder:480b CloudOllamafreefree262K
9AutoOpenRouterfreefree2M
10FreeOpenRouterfreefree200K
11Nemotron 3.5 Lightning:freeOpenRouterfreefree1M
12Laguna S 2.1:freeOpenRouterfreefree262K
13Laguna Xs 2.1:freeOpenRouterfreefree262K
14Glm 5.2:freeOpenRouterfreefree256K
15Nemotron 3 Ultra 550b A55b:freeOpenRouterfreefree1M
16Minimax M3:freeOpenRouterfreefree1.0M
17Nemotron 3 Nano Omni 30b A3b Reasoning:freeOpenRouterfreefree256K
18Gemma 4 26b A4b It:freeOpenRouterfreefree262K
19Gemma 4 31b It:freeOpenRouterfreefree262K
20Nemotron 3 Super 120b A12b:freeOpenRouterfreefree262K
21Ternary Bonsai 27BTogether AIfreefree262K
22Qwen3.7 FlashOpenRouter$0.03$0.13$0.0061M
23Gemma 4 26b A4b ItOpenRouter$0.042$0.22262K
24Databricks GPT 5 NanoDatabricks$0.05$0.4$0.005272K
25Qwen TurboAlibaba DashScope$0.05$0.21M
26Qwen TurboAlibaba DashScope$0.05$0.21M
27Qwen TurboAlibaba DashScope$0.05$0.21M
28GPT 5 NanoOpenAI$0.05$0.4$0.005272K
29GPT 5 NanoOpenAI$0.05$0.4$0.005272K
30GPT 5 NanoAzure OpenAI$0.05$0.4$0.005272K
31Nemotron 3 Nano 30B A3BDeepInfra$0.05$0.2$0.025262K
32Nemotron Lightning 3p5 30b A3bFireworks AI$0.05$0.2$0.01262K
33Nemotron 3 Nano 30b A3bNovita AI$0.05$0.2262K
34GPT 5 NanoOpenRouter$0.05$0.4$0.005272K
35Nemotron 3 Nano 30b A3bOpenRouter$0.05$0.2$0.03262K
36GPT 5 NanoAzure OpenAI$0.055$0.44$0.0055272K
37GLM 4.7 FlashDeepInfra$0.06$0.4$0.01203K
38NVIDIA Nemotron 3 Nano 30B A3BNebius$0.06$0.24262K
39Nemotron 3 Nano OmniNebius$0.06$0.24262K
40Nemotron 3 5 LightningNebius$0.06$0.241.0M
41Ling 3.0 Flash FastNovita AI$0.06$0.18$0.012262K
42Ling 3.0 FlashNovita AI$0.06$0.18$0.012262K
43Glm 4.7 FlashOpenRouter$0.06$0.4$0.01200K
44Laguna Xs 2.1OpenRouter$0.06$0.12$0.03262K
45Qwen3.5 Flash 02 23OpenRouter$0.065$0.261M
46Deepseek V4 Flash 0731OpenRouter$0.065$0.18$0.0161.3M
47Gemma 4 26B A4B ItDeepInfra$0.07$0.34262K
48Glm 4.7 FlashNovita AI$0.07$0.4$0.01200K
49Qwen3 Coder 30b A3b InstructOpenRouter$0.07$0.28262K
50Amazon.nova Lite V1:0Amazon Bedrock$0.072$0.288300K
51Nvidia.nemotron Nano 3 30bAmazon Bedrock$0.072$0.288262K
52Gemini 2.0 Flash LiteGoogle Gemini$0.075$0.3$0.0191.0M
53Gemini 2.0 Flash Lite 001Google Gemini$0.075$0.3$0.0191.0M
54NVIDIA Nemotron 3.5 LightningDeepInfra$0.08$0.2$0.04262K
55DeepSeek V4 Flash 0731DeepInfra$0.08$0.18$0.0161.0M
56Nemotron 3.5 LightningOpenRouter$0.08$0.2262K
57NVIDIA Nemotron 3 Super 120B A12BDeepInfra$0.085$0.4262K
58Nemotron 3 Super 120b A12bOpenRouter$0.085$0.41M
59Deepseek V4 FlashOpenRouter$0.085$0.171$0.0171.0M
60Qwen3 235B A22B Instruct 2507DeepInfra$0.09$0.55262K

Other use cases

Frequently asked

Which AI model is cheapest for very long documents?

Gemini Exp 1114 from Google Gemini currently leads this list, at $0/1M input tokens. Rankings are recomputed from the raw dataset on every refresh, so they track price cuts automatically.

How are these rankings calculated?

Each use case applies a capability filter (context size, vision, tool calling, cache pricing) and then ranks the surviving models by the blended cost that profile actually pays. See the methodology page for the exact formulas.

Is the cheapest model always the right choice?

No. Check the capability flags and context window, and benchmark quality on your own data — the cheapest model that fails your evals costs more than the second cheapest that passes them.