Cheapest model for summarisation at scale
Ranked for an input-heavy workload (long documents, short summaries), where output tokens matter far less than input.
Current leader: Gemini Exp 1114 (Google Gemini). Refreshed 2026-09-13 13:40 UTC.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini Exp 1114 | Google Gemini | free | free | — | 1.0M |
| 2 | Gemini Exp 1206 | Google Gemini | free | free | — | 2.1M |
| 3 | Gemma 3 27b It | Google Gemini | free | free | — | 131K |
| 4 | Gemma 4 26b A4b It | Google Gemini | free | free | — | 262K |
| 5 | Gemma 4 31b It | Google Gemini | free | free | — | 262K |
| 6 | Labs Leanstral 1 5 | Mistral AI | free | free | — | 262K |
| 7 | Labs Leanstral 1 5 1 | Mistral AI | free | free | — | 262K |
| 8 | Anthropic.claude Mythos Preview | Amazon Bedrock | free | free | — | 1M |
| 9 | Deepseek V3.1:671b Cloud | Ollama | free | free | — | 164K |
| 10 | GPT Oss:120b Cloud | Ollama | free | free | — | 131K |
| 11 | GPT Oss:20b Cloud | Ollama | free | free | — | 131K |
| 12 | Qwen3 Coder:480b Cloud | Ollama | free | free | — | 262K |
| 13 | Auto | OpenRouter | free | free | — | 2M |
| 14 | Free | OpenRouter | free | free | — | 200K |
| 15 | Bodybuilder | OpenRouter | free | free | — | 128K |
| 16 | Nemotron 3.5 Lightning:free | OpenRouter | free | free | — | 1M |
| 17 | Laguna S 2.1:free | OpenRouter | free | free | — | 262K |
| 18 | Laguna Xs 2.1:free | OpenRouter | free | free | — | 262K |
| 19 | Glm 5.2:free | OpenRouter | free | free | — | 256K |
| 20 | Nemotron 3.5 Content Safety:free | OpenRouter | free | free | — | 128K |
| 21 | Nemotron 3 Ultra 550b A55b:free | OpenRouter | free | free | — | 1M |
| 22 | Minimax M3:free | OpenRouter | free | free | — | 1.0M |
| 23 | Nemotron 3 Nano Omni 30b A3b Reasoning:free | OpenRouter | free | free | — | 256K |
| 24 | Gemma 4 26b A4b It:free | OpenRouter | free | free | — | 262K |
| 25 | Gemma 4 31b It:free | OpenRouter | free | free | — | 262K |
| 26 | Minimax M2.7:free | OpenRouter | free | free | — | 197K |
| 27 | Nemotron 3 Super 120b A12b:free | OpenRouter | free | free | — | 262K |
| 28 | Ternary Bonsai 27B | Together AI | free | free | — | 262K |
| 29 | Llama 3.2 3B Instruct | DeepInfra | $0.02 | $0.02 | — | 131K |
| 30 | Llama 3.2 1b Instruct | Novita AI | $0.02 | $0.02 | — | 131K |
| 31 | Mistral Nemo Instruct 2407 | DeepInfra | $0.019 | $0.03 | — | 131K |
| 32 | Mistral Nemo | OpenRouter | $0.019 | $0.03 | — | 131K |
| 33 | Meta Llama 3.1 8B Instruct Turbo | DeepInfra | $0.02 | $0.04 | — | 131K |
| 34 | Llama Guard 3 8B | Nebius | $0.02 | $0.06 | — | 128K |
| 35 | Meta Llama 3.1 8B Instruct | Nebius | $0.02 | $0.06 | — | 128K |
| 36 | Qwen2 VL 7B Instruct | Nebius | $0.02 | $0.06 | — | 131K |
| 37 | Granite 4.0 H Micro | Cloudflare Workers AI | $0.017 | $0.112 | — | 131K |
| 38 | Gemma 4 E4B It | DeepInfra | $0.02 | $0.1 | — | 131K |
| 39 | Qwen3 4b Fp8 | Novita AI | $0.03 | $0.03 | — | 128K |
| 40 | Meta Llama 3.1 8B Instruct | DeepInfra | $0.03 | $0.05 | — | 131K |
| 41 | GPT Oss 20b | OpenRouter | $0.03 | $0.13 | — | 131K |
| 42 | Qwen3.7 Flash | OpenRouter | $0.03 | $0.13 | $0.006 | 1M |
| 43 | GPT Oss 20b | DeepInfra | $0.03 | $0.14 | — | 131K |
| 44 | Qwen3 8b Fp8 | Novita AI | $0.035 | $0.138 | — | 128K |
| 45 | Mistral Nemo Instruct 2407 | Nebius | $0.04 | $0.12 | — | 128K |
| 46 | Llama 3.2 11B Vision Instruct | DeepInfra | $0.049 | $0.049 | — | 131K |
| 47 | GPT Oss 120b | DeepInfra | $0.037 | $0.17 | — | 131K |
| 48 | GPT Oss 120b | OpenRouter | $0.037 | $0.17 | — | 131K |
| 49 | GPT Oss 20b | Novita AI | $0.04 | $0.15 | — | 131K |
| 50 | NVIDIA Nemotron Nano 9B V2 | DeepInfra | $0.04 | $0.16 | — | 131K |
| 51 | Llama 3.1 8b Instant | Groq | $0.05 | $0.08 | — | 131K |
| 52 | Llama 3.1 8b Instruct | OpenRouter | $0.05 | $0.08 | $0.025 | 131K |
| 53 | Llama Guard 3 8B | DeepInfra | $0.055 | $0.055 | — | 131K |
| 54 | Gemma 3 4b It | DeepInfra | $0.05 | $0.1 | — | 131K |
| 55 | Gemma 3 12b It | Novita AI | $0.05 | $0.1 | — | 131K |
| 56 | Gemma 3 4b It | OpenRouter | $0.05 | $0.1 | — | 131K |
| 57 | Amazon.nova Micro V1:0 | Amazon Bedrock | $0.042 | $0.168 | — | 128K |
| 58 | Gemma 3 12b It | DeepInfra | $0.05 | $0.15 | — | 131K |
| 59 | Gemma 3 12b It | OpenRouter | $0.05 | $0.15 | — | 131K |
| 60 | Gemma 4 26b A4b It | OpenRouter | $0.042 | $0.22 | — | 262K |
Other use cases
Frequently asked
› What is the cheapest model for summarising thousands of documents?
Gemini Exp 1114 from Google Gemini currently leads this list, at $0/1M input tokens. Rankings are recomputed from the raw dataset on every refresh, so they track price cuts automatically.
› How are these rankings calculated?
Each use case applies a capability filter (context size, vision, tool calling, cache pricing) and then ranks the surviving models by the blended cost that profile actually pays. See the methodology page for the exact formulas.
› Is the cheapest model always the right choice?
No. Check the capability flags and context window, and benchmark quality on your own data — the cheapest model that fails your evals costs more than the second cheapest that passes them.