Price per 1M logo

Cheapest model for summarisation at scale

Ranked for an input-heavy workload (long documents, short summaries), where output tokens matter far less than input.

Current leader: Gemini Exp 1114 (Google Gemini). Refreshed 2026-09-13 13:40 UTC.

#ModelProviderInput $/1MOutput $/1MCached $/1MContext
1Gemini Exp 1114Google Geminifreefree1.0M
2Gemini Exp 1206Google Geminifreefree2.1M
3Gemma 3 27b ItGoogle Geminifreefree131K
4Gemma 4 26b A4b ItGoogle Geminifreefree262K
5Gemma 4 31b ItGoogle Geminifreefree262K
6Labs Leanstral 1 5Mistral AIfreefree262K
7Labs Leanstral 1 5 1Mistral AIfreefree262K
8Anthropic.claude Mythos PreviewAmazon Bedrockfreefree1M
9Deepseek V3.1:671b CloudOllamafreefree164K
10GPT Oss:120b CloudOllamafreefree131K
11GPT Oss:20b CloudOllamafreefree131K
12Qwen3 Coder:480b CloudOllamafreefree262K
13AutoOpenRouterfreefree2M
14FreeOpenRouterfreefree200K
15BodybuilderOpenRouterfreefree128K
16Nemotron 3.5 Lightning:freeOpenRouterfreefree1M
17Laguna S 2.1:freeOpenRouterfreefree262K
18Laguna Xs 2.1:freeOpenRouterfreefree262K
19Glm 5.2:freeOpenRouterfreefree256K
20Nemotron 3.5 Content Safety:freeOpenRouterfreefree128K
21Nemotron 3 Ultra 550b A55b:freeOpenRouterfreefree1M
22Minimax M3:freeOpenRouterfreefree1.0M
23Nemotron 3 Nano Omni 30b A3b Reasoning:freeOpenRouterfreefree256K
24Gemma 4 26b A4b It:freeOpenRouterfreefree262K
25Gemma 4 31b It:freeOpenRouterfreefree262K
26Minimax M2.7:freeOpenRouterfreefree197K
27Nemotron 3 Super 120b A12b:freeOpenRouterfreefree262K
28Ternary Bonsai 27BTogether AIfreefree262K
29Llama 3.2 3B InstructDeepInfra$0.02$0.02131K
30Llama 3.2 1b InstructNovita AI$0.02$0.02131K
31Mistral Nemo Instruct 2407DeepInfra$0.019$0.03131K
32Mistral NemoOpenRouter$0.019$0.03131K
33Meta Llama 3.1 8B Instruct TurboDeepInfra$0.02$0.04131K
34Llama Guard 3 8BNebius$0.02$0.06128K
35Meta Llama 3.1 8B InstructNebius$0.02$0.06128K
36Qwen2 VL 7B InstructNebius$0.02$0.06131K
37Granite 4.0 H MicroCloudflare Workers AI$0.017$0.112131K
38Gemma 4 E4B ItDeepInfra$0.02$0.1131K
39Qwen3 4b Fp8Novita AI$0.03$0.03128K
40Meta Llama 3.1 8B InstructDeepInfra$0.03$0.05131K
41GPT Oss 20bOpenRouter$0.03$0.13131K
42Qwen3.7 FlashOpenRouter$0.03$0.13$0.0061M
43GPT Oss 20bDeepInfra$0.03$0.14131K
44Qwen3 8b Fp8Novita AI$0.035$0.138128K
45Mistral Nemo Instruct 2407Nebius$0.04$0.12128K
46Llama 3.2 11B Vision InstructDeepInfra$0.049$0.049131K
47GPT Oss 120bDeepInfra$0.037$0.17131K
48GPT Oss 120bOpenRouter$0.037$0.17131K
49GPT Oss 20bNovita AI$0.04$0.15131K
50NVIDIA Nemotron Nano 9B V2DeepInfra$0.04$0.16131K
51Llama 3.1 8b InstantGroq$0.05$0.08131K
52Llama 3.1 8b InstructOpenRouter$0.05$0.08$0.025131K
53Llama Guard 3 8BDeepInfra$0.055$0.055131K
54Gemma 3 4b ItDeepInfra$0.05$0.1131K
55Gemma 3 12b ItNovita AI$0.05$0.1131K
56Gemma 3 4b ItOpenRouter$0.05$0.1131K
57Amazon.nova Micro V1:0Amazon Bedrock$0.042$0.168128K
58Gemma 3 12b ItDeepInfra$0.05$0.15131K
59Gemma 3 12b ItOpenRouter$0.05$0.15131K
60Gemma 4 26b A4b ItOpenRouter$0.042$0.22262K

Other use cases

Frequently asked

What is the cheapest model for summarising thousands of documents?

Gemini Exp 1114 from Google Gemini currently leads this list, at $0/1M input tokens. Rankings are recomputed from the raw dataset on every refresh, so they track price cuts automatically.

How are these rankings calculated?

Each use case applies a capability filter (context size, vision, tool calling, cache pricing) and then ranks the surviving models by the blended cost that profile actually pays. See the methodology page for the exact formulas.

Is the cheapest model always the right choice?

No. Check the capability flags and context window, and benchmark quality on your own data — the cheapest model that fails your evals costs more than the second cheapest that passes them.