Cheapest long-context models
Models with at least a 200K-token context window, ranked by input price — for summarising books, legal dumps and codebases.
Current leader: Gemini Exp 1114 (Google Gemini). Refreshed 2026-09-13 13:40 UTC.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini Exp 1114 | Google Gemini | free | free | — | 1.0M |
| 2 | Gemini Exp 1206 | Google Gemini | free | free | — | 2.1M |
| 3 | Gemma 4 26b A4b It | Google Gemini | free | free | — | 262K |
| 4 | Gemma 4 31b It | Google Gemini | free | free | — | 262K |
| 5 | Labs Leanstral 1 5 | Mistral AI | free | free | — | 262K |
| 6 | Labs Leanstral 1 5 1 | Mistral AI | free | free | — | 262K |
| 7 | Anthropic.claude Mythos Preview | Amazon Bedrock | free | free | — | 1M |
| 8 | Qwen3 Coder:480b Cloud | Ollama | free | free | — | 262K |
| 9 | Auto | OpenRouter | free | free | — | 2M |
| 10 | Free | OpenRouter | free | free | — | 200K |
| 11 | Nemotron 3.5 Lightning:free | OpenRouter | free | free | — | 1M |
| 12 | Laguna S 2.1:free | OpenRouter | free | free | — | 262K |
| 13 | Laguna Xs 2.1:free | OpenRouter | free | free | — | 262K |
| 14 | Glm 5.2:free | OpenRouter | free | free | — | 256K |
| 15 | Nemotron 3 Ultra 550b A55b:free | OpenRouter | free | free | — | 1M |
| 16 | Minimax M3:free | OpenRouter | free | free | — | 1.0M |
| 17 | Nemotron 3 Nano Omni 30b A3b Reasoning:free | OpenRouter | free | free | — | 256K |
| 18 | Gemma 4 26b A4b It:free | OpenRouter | free | free | — | 262K |
| 19 | Gemma 4 31b It:free | OpenRouter | free | free | — | 262K |
| 20 | Nemotron 3 Super 120b A12b:free | OpenRouter | free | free | — | 262K |
| 21 | Ternary Bonsai 27B | Together AI | free | free | — | 262K |
| 22 | Qwen3.7 Flash | OpenRouter | $0.03 | $0.13 | $0.006 | 1M |
| 23 | Gemma 4 26b A4b It | OpenRouter | $0.042 | $0.22 | — | 262K |
| 24 | Databricks GPT 5 Nano | Databricks | $0.05 | $0.4 | $0.005 | 272K |
| 25 | Qwen Turbo | Alibaba DashScope | $0.05 | $0.2 | — | 1M |
| 26 | Qwen Turbo | Alibaba DashScope | $0.05 | $0.2 | — | 1M |
| 27 | Qwen Turbo | Alibaba DashScope | $0.05 | $0.2 | — | 1M |
| 28 | GPT 5 Nano | OpenAI | $0.05 | $0.4 | $0.005 | 272K |
| 29 | GPT 5 Nano | OpenAI | $0.05 | $0.4 | $0.005 | 272K |
| 30 | GPT 5 Nano | Azure OpenAI | $0.05 | $0.4 | $0.005 | 272K |
| 31 | Nemotron 3 Nano 30B A3B | DeepInfra | $0.05 | $0.2 | $0.025 | 262K |
| 32 | Nemotron Lightning 3p5 30b A3b | Fireworks AI | $0.05 | $0.2 | $0.01 | 262K |
| 33 | Nemotron 3 Nano 30b A3b | Novita AI | $0.05 | $0.2 | — | 262K |
| 34 | GPT 5 Nano | OpenRouter | $0.05 | $0.4 | $0.005 | 272K |
| 35 | Nemotron 3 Nano 30b A3b | OpenRouter | $0.05 | $0.2 | $0.03 | 262K |
| 36 | GPT 5 Nano | Azure OpenAI | $0.055 | $0.44 | $0.0055 | 272K |
| 37 | GLM 4.7 Flash | DeepInfra | $0.06 | $0.4 | $0.01 | 203K |
| 38 | NVIDIA Nemotron 3 Nano 30B A3B | Nebius | $0.06 | $0.24 | — | 262K |
| 39 | Nemotron 3 Nano Omni | Nebius | $0.06 | $0.24 | — | 262K |
| 40 | Nemotron 3 5 Lightning | Nebius | $0.06 | $0.24 | — | 1.0M |
| 41 | Ling 3.0 Flash Fast | Novita AI | $0.06 | $0.18 | $0.012 | 262K |
| 42 | Ling 3.0 Flash | Novita AI | $0.06 | $0.18 | $0.012 | 262K |
| 43 | Glm 4.7 Flash | OpenRouter | $0.06 | $0.4 | $0.01 | 200K |
| 44 | Laguna Xs 2.1 | OpenRouter | $0.06 | $0.12 | $0.03 | 262K |
| 45 | Qwen3.5 Flash 02 23 | OpenRouter | $0.065 | $0.26 | — | 1M |
| 46 | Deepseek V4 Flash 0731 | OpenRouter | $0.065 | $0.18 | $0.016 | 1.3M |
| 47 | Gemma 4 26B A4B It | DeepInfra | $0.07 | $0.34 | — | 262K |
| 48 | Glm 4.7 Flash | Novita AI | $0.07 | $0.4 | $0.01 | 200K |
| 49 | Qwen3 Coder 30b A3b Instruct | OpenRouter | $0.07 | $0.28 | — | 262K |
| 50 | Amazon.nova Lite V1:0 | Amazon Bedrock | $0.072 | $0.288 | — | 300K |
| 51 | Nvidia.nemotron Nano 3 30b | Amazon Bedrock | $0.072 | $0.288 | — | 262K |
| 52 | Gemini 2.0 Flash Lite | Google Gemini | $0.075 | $0.3 | $0.019 | 1.0M |
| 53 | Gemini 2.0 Flash Lite 001 | Google Gemini | $0.075 | $0.3 | $0.019 | 1.0M |
| 54 | NVIDIA Nemotron 3.5 Lightning | DeepInfra | $0.08 | $0.2 | $0.04 | 262K |
| 55 | DeepSeek V4 Flash 0731 | DeepInfra | $0.08 | $0.18 | $0.016 | 1.0M |
| 56 | Nemotron 3.5 Lightning | OpenRouter | $0.08 | $0.2 | — | 262K |
| 57 | NVIDIA Nemotron 3 Super 120B A12B | DeepInfra | $0.085 | $0.4 | — | 262K |
| 58 | Nemotron 3 Super 120b A12b | OpenRouter | $0.085 | $0.4 | — | 1M |
| 59 | Deepseek V4 Flash | OpenRouter | $0.085 | $0.171 | $0.017 | 1.0M |
| 60 | Qwen3 235B A22B Instruct 2507 | DeepInfra | $0.09 | $0.55 | — | 262K |
Other use cases
Frequently asked
› Which AI model is cheapest for very long documents?
Gemini Exp 1114 from Google Gemini currently leads this list, at $0/1M input tokens. Rankings are recomputed from the raw dataset on every refresh, so they track price cuts automatically.
› How are these rankings calculated?
Each use case applies a capability filter (context size, vision, tool calling, cache pricing) and then ranks the surviving models by the blended cost that profile actually pays. See the methodology page for the exact formulas.
› Is the cheapest model always the right choice?
No. Check the capability flags and context window, and benchmark quality on your own data — the cheapest model that fails your evals costs more than the second cheapest that passes them.