AI model price report — week 39, 2026
93 price movements across 5 providers. This report is generated from the same dataset that powers the JSON API and the raw changelog — 2,045 priced chat models across 33 providers. Dataset refreshed 2026-09-21 00:12 UTC.
Price movements this week
| Model | Provider | Field | Before | After | Change |
|---|---|---|---|---|---|
| Gemini Pro | Google Gemini | input per 1m | $1.25 | $2 | +60.0% |
| Gemini Pro | Google Gemini | output per 1m | $10 | $12 | +20.0% |
| Gemini Pro | Google Gemini | cached input per 1m | $0.125 | $0.2 | +60.0% |
| Gemini Robotics Er 2 Preview | Google Gemini | input per 1m | $2 | $1 | -50.0% |
| Gemini Robotics Er 2 Preview | Google Gemini | output per 1m | $10 | $5 | -50.0% |
| Gemini Robotics Er 2 Preview | Google Gemini | cached input per 1m | $0.2 | $0.1 | -50.0% |
| Gemini Flash | Google Gemini | input per 1m | $0.3 | $0.75 | +150.0% |
| Gemini Flash | Google Gemini | output per 1m | $2.5 | $3.75 | +50.0% |
| Gemini Flash | Google Gemini | cached input per 1m | $0.03 | $0.075 | +150.0% |
| Gemini Flash Lite | Google Gemini | input per 1m | $0.1 | $0.3 | +200.0% |
| Gemini Flash Lite | Google Gemini | output per 1m | $0.4 | $2.5 | +525.0% |
| Gemini Flash Lite | Google Gemini | cached input per 1m | $0.01 | $0.03 | +200.0% |
| GPT 5.6 Sol | Azure OpenAI | input per 1m | $5 | $4 | -20.0% |
| GPT 5.6 Sol | Azure OpenAI | output per 1m | $30 | $20 | -33.3% |
| GPT 5.6 Sol | Azure OpenAI | cached input per 1m | $0.5 | $0.4 | -20.0% |
| GPT 5.1 | Azure OpenAI | input per 1m | $1.38 | $1.38 | -0.4% |
| GPT 5.1 | Azure OpenAI | cached input per 1m | $0.14 | $0.138 | -1.8% |
| GPT 5.1 Chat | Azure OpenAI | input per 1m | $1.38 | $1.38 | -0.4% |
| GPT 5.1 Chat | Azure OpenAI | cached input per 1m | $0.14 | $0.138 | -1.8% |
| GPT 5.1 Codex | Azure OpenAI | input per 1m | $1.38 | $1.38 | -0.4% |
| GPT 5.1 Codex | Azure OpenAI | cached input per 1m | $0.14 | $0.138 | -1.8% |
| O1 Mini | Azure OpenAI | input per 1m | $1.21 | $1.1 | -9.1% |
| O1 Mini | Azure OpenAI | output per 1m | $4.84 | $4.4 | -9.1% |
| O1 Mini | Azure OpenAI | cached input per 1m | $0.605 | $0.55 | -9.1% |
| GPT 5.1 Codex Mini | Azure OpenAI | cached input per 1m | $0.028 | $0.028 | -1.8% |
| Deepseek V4 Pro | Fireworks AI | input per 1m | $1.74 | $1.2 | -31.0% |
| Deepseek V4 Pro | Fireworks AI | output per 1m | $3.48 | $1.2 | -65.5% |
| Deepseek V4 Pro | Fireworks AI | cached input per 1m | $0.145 | $0.6 | +313.8% |
| Kimi K3 | OpenRouter | input per 1m | $2.1 | $1.7 | -19.1% |
| Kimi K3 | OpenRouter | output per 1m | $10.53 | $8.5 | -19.3% |
| Kimi K3 | OpenRouter | cached input per 1m | $0.235 | $0.17 | -27.7% |
| Deepseek V4 Pro 0813 | OpenRouter | input per 1m | $0.579 | $1.32 | +127.8% |
| Deepseek V4 Pro 0813 | OpenRouter | output per 1m | $1.74 | $3.96 | +127.8% |
| Deepseek V4 Pro 0813 | OpenRouter | cached input per 1m | $0.019 | $0.044 | +127.8% |
| Glm 5.3 | OpenRouter | input per 1m | $1.4 | $0.91 | -35.0% |
| Glm 5.3 | OpenRouter | output per 1m | $4.4 | $2.86 | -35.0% |
| Glm 5.3 | OpenRouter | cached input per 1m | $0.14 | $0.169 | +20.7% |
| Kimi K2.7 Code | OpenRouter | input per 1m | $0.71 | $0.706 | -0.5% |
| Kimi K2.7 Code | OpenRouter | output per 1m | $3.5 | $3.21 | -8.3% |
| Kimi K2.7 Code | OpenRouter | cached input per 1m | $0.15 | $0.18 | +20.0% |
| Glm 5.2 | OpenRouter | input per 1m | $0.6 | $0.65 | +8.3% |
| Glm 5.2 | OpenRouter | output per 1m | $2 | $2.04 | +2.1% |
| Glm 5.2 | OpenRouter | cached input per 1m | $0.15 | $0.121 | -19.6% |
| Nemotron 3 Ultra 550b A55b | OpenRouter | input per 1m | $0.625 | $0.6 | -4.0% |
| Nemotron 3 Ultra 550b A55b | OpenRouter | output per 1m | $3.13 | $2.4 | -23.2% |
| Nemotron 3 Ultra 550b A55b | OpenRouter | cached input per 1m | $0.188 | $0.12 | -36.0% |
| Mistral Large 2512 | OpenRouter | input per 1m | $0.5 | $0.55 | +10.0% |
| Mistral Large 2512 | OpenRouter | output per 1m | $1.5 | $1.65 | +10.0% |
| Deepseek V4 Pro | OpenRouter | input per 1m | $0.86 | $0.422 | -50.9% |
| Deepseek V4 Pro | OpenRouter | output per 1m | $1.72 | $0.845 | -50.9% |
| Deepseek V4 Pro | OpenRouter | cached input per 1m | $0.072 | $0.035 | -50.9% |
| Minimax M1 | OpenRouter | input per 1m | $0.55 | $0.4 | -27.3% |
| Remm Slerp L2 13b | OpenRouter | input per 1m | $0.45 | $0.35 | -22.2% |
| Deepseek V4.1 Flash | OpenRouter | input per 1m | $0.15 | $0.3 | +100.0% |
| Deepseek V4.1 Flash | OpenRouter | output per 1m | $0.6 | $1.2 | +100.0% |
| Deepseek V4.1 Flash | OpenRouter | cached input per 1m | $0.003 | $0.006 | +100.0% |
| Qwen3.8 27b | OpenRouter | input per 1m | $0.42 | $0.2 | -52.4% |
| Qwen3.8 27b | OpenRouter | output per 1m | $3 | $2.55 | -15.0% |
| Llama 4 Maverick | OpenRouter | input per 1m | $0.2 | $0.188 | -6.3% |
| Llama 4 Maverick | OpenRouter | output per 1m | $0.696 | $0.652 | -6.3% |
| GPT Oss 120b | OpenRouter | input per 1m | $0.037 | $0.15 | +305.4% |
| GPT Oss 120b | OpenRouter | output per 1m | $0.17 | $0.6 | +252.9% |
| Qwen3.6 35b A3b | OpenRouter | input per 1m | $0.1 | $0.15 | +50.0% |
| Qwen3.6 35b A3b | OpenRouter | output per 1m | $0.9 | $1 | +11.1% |
| Qwen3 Vl 30b A3b Instruct | OpenRouter | input per 1m | $0.15 | $0.13 | -13.3% |
| Qwen3 Vl 30b A3b Instruct | OpenRouter | output per 1m | $0.6 | $0.52 | -13.3% |
| Qwen3 14b | OpenRouter | input per 1m | $0.228 | $0.12 | -47.3% |
| Qwen3 14b | OpenRouter | output per 1m | $0.91 | $0.24 | -73.6% |
| Mistral Small 3.2 24b Instruct | OpenRouter | input per 1m | $0.075 | $0.094 | +25.0% |
| Mistral Small 3.2 24b Instruct | OpenRouter | output per 1m | $0.2 | $0.25 | +25.0% |
| Glm 5.3 Flash | OpenRouter | input per 1m | $0.15 | $0.09 | -40.0% |
| Glm 5.3 Flash | OpenRouter | output per 1m | $0.5 | $0.3 | -40.0% |
| Glm 5.3 Flash | OpenRouter | cached input per 1m | $0.03 | $0.018 | -40.0% |
| Gemma 4 26b A4b It | OpenRouter | input per 1m | $0.042 | $0.09 | +114.3% |
| Gemma 4 26b A4b It | OpenRouter | output per 1m | $0.22 | $0.3 | +36.4% |
| Qwen3 235b A22b 2507 | OpenRouter | input per 1m | $0.22 | $0.087 | -60.2% |
| Qwen3 235b A22b 2507 | OpenRouter | output per 1m | $0.88 | $0.35 | -60.2% |
| Mythomax L2 13b | OpenRouter | input per 1m | $0.06 | $0.08 | +33.3% |
| Mythomax L2 13b | OpenRouter | output per 1m | $0.06 | $0.11 | +83.3% |
| Nemotron 3 Super 120b A12b | OpenRouter | input per 1m | $0.085 | $0.08 | -5.9% |
| Nemotron 3 Super 120b A12b | OpenRouter | output per 1m | $0.4 | $0.45 | +12.5% |
| Nemotron 3.5 Lightning | OpenRouter | input per 1m | $0.08 | $0.07 | -12.5% |
| Glm 4.7 Flash | OpenRouter | input per 1m | $0.06 | $0.06 | +0.8% |
| Nemotron 3 Nano 30b A3b | OpenRouter | input per 1m | $0.05 | $0.06 | +20.0% |
| Nemotron 3 Nano 30b A3b | OpenRouter | output per 1m | $0.2 | $0.24 | +20.0% |
| Qwen3 30b A3b Instruct 2507 | OpenRouter | input per 1m | $0.09 | $0.048 | -46.5% |
| Qwen3 30b A3b Instruct 2507 | OpenRouter | output per 1m | $0.3 | $0.193 | -35.6% |
| Deepseek V4 Flash 0731 | OpenRouter | input per 1m | $0.065 | $0.04 | -38.5% |
| Deepseek V4 Flash 0731 | OpenRouter | output per 1m | $0.18 | $0.16 | -11.1% |
| Deepseek V4 Flash | OpenRouter | input per 1m | $0.085 | $0.036 | -58.4% |
| Deepseek V4 Flash | OpenRouter | output per 1m | $0.171 | $0.071 | -58.4% |
| Deepseek V4 Flash | OpenRouter | cached input per 1m | $0.017 | $0.0071 | -58.4% |
| Gemini 2.5 Flash | Replicate | input per 1m | $2.5 | $0.3 | -88.0% |
- Gemini 2.5 Pro Preview Tts-99.2%
- Qwen3 Next 80b A3b Instruct-93.1%
- Gemini 2.5 Flash-88.0%
- Glm 4.6-87.5%
- Mistral Small 3.2 24b Instruct-87.2%
- Deepseek Chat V3 0324+1700.0%
- Gemma 4 26b A4b It+1340.0%
- Nemotron 3 Super 120b A12b+1340.0%
- Mistral Large+1150.2%
- Gemini Omni 1.1 Flash+700.0%
Models added
- Qwen3.8 FlashAlibaba DashScope
- Qwen3.8 Omni FlashAlibaba DashScope
- Zai Glm 5 3Mistral AI
- Zai Glm 5Mistral AI
- Zai GlmMistral AI
- GPT 5.5 CyberOpenAI
- GPT Rosalind ResearchOpenAI
- GPT 6 AstraAzure OpenAI
- GPT ChatAzure OpenAI
- ChatAzure OpenAI
- GPT 5.5Azure OpenAI
- GPT 5.6 SolAzure OpenAI
Models retired
No tracked endpoints disappeared from the upstream sources this week.
Cheapest input tokens
The floor of the market as of this data refresh — what the input side of a bill costs per 1M tokens.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Qwen2.5 Coder 7B | Nebius | $0.01 | $0.03 | — | 33K |
| 2 | Qwen2.5 Coder 3B Instruct | Nscale | $0.01 | $0.03 | — | — |
| 3 | Qwen2.5 Coder 7B Instruct | Nscale | $0.01 | $0.03 | — | — |
| 4 | Nemotron 3.5 Lightning 30b A3b | Perplexity | $0.011 | $0.17 | $0.0011 | — |
| 5 | Granite 4.0 H Micro | Cloudflare Workers AI | $0.017 | $0.112 | — | 131K |
| 6 | Granite 4.0 H Micro | OpenRouter | $0.017 | $0.112 | — | 131K |
| 7 | Mistral Nemo Instruct 2407 | DeepInfra | $0.019 | $0.03 | — | 131K |
| 8 | Mistral Nemo | OpenRouter | $0.019 | $0.03 | — | 131K |
| 9 | Llama 3.2 3B Instruct | DeepInfra | $0.02 | $0.02 | — | 131K |
| 10 | Meta Llama 3.1 8B Instruct Turbo | DeepInfra | $0.02 | $0.04 | — | 131K |
| 11 | Gemma 4 E4B It | DeepInfra | $0.02 | $0.1 | — | 131K |
| 12 | Llama Guard 3 8B | Nebius | $0.02 | $0.06 | — | 128K |
Best blended cost
Weighted 3:1 toward input tokens, which is how most chat and agent workloads actually bill.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini Exp 1114 | Google Gemini | free | free | — | 1.0M |
| 2 | Gemini Exp 1206 | Google Gemini | free | free | — | 2.1M |
| 3 | Gemma 3 27b It | Google Gemini | free | free | — | 131K |
| 4 | Gemma 4 26b A4b It | Google Gemini | free | free | — | 262K |
| 5 | Gemma 4 31b It | Google Gemini | free | free | — | 262K |
| 6 | Learnlm 1.5 Pro Experimental | Google Gemini | free | free | — | 33K |
| 7 | Labs Leanstral 1 5 | Mistral AI | free | free | — | 262K |
| 8 | Labs Leanstral 1 5 1 | Mistral AI | free | free | — | 262K |
| 9 | Anthropic.claude Mythos Preview | Amazon Bedrock | free | free | — | 1M |
| 10 | Gemma 2b It Lora | Cloudflare Workers AI | free | free | — | 8K |
| 11 | Mistral 7b Instruct V0.2 Lora | Cloudflare Workers AI | free | free | — | 15K |
| 12 | Llama 2 7b Chat Hf Lora | Cloudflare Workers AI | free | free | — | 8K |
Most expensive models tracked
The top of the market, for reference — what frontier pricing looks like against the floor above.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | O1 Pro | OpenRouter | $150 | $600 | — | 200K |
| 2 | O1 Pro | OpenAI | $150 | $600 | — | 200K |
| 3 | O1 Pro | OpenAI | $150 | $600 | — | 200K |
| 4 | GPT 4.5 Preview | Azure OpenAI | $75 | $150 | $37.5 | 128K |
| 5 | GPT 4 32k 0613 | Azure OpenAI | $60 | $120 | — | 33K |
| 6 | GPT 4 32k | Azure OpenAI | $60 | $120 | — | 33K |
| 7 | GPT 5.4 Pro | OpenRouter | $30 | $180 | — | 1.1M |
| 8 | GPT 5.5 Pro | OpenRouter | $30 | $180 | — | 1.1M |
What this costs in practice
Three common workloads priced at the cheapest rate available in the index this week.
| Workload | Priced at | Total |
|---|---|---|
| support bot · 10,000 chats/day for a month (1.2k in, 400 out) | Qwen2.5 Coder 7B · $0.01/$0.03 | $7.20 |
| 1,000,000 document summaries (8k in, 300 out) | Qwen2.5 Coder 7B · $0.01/$0.03 | $89.00 |
| RAG index · 100M embedding tokens | Pplx Embed V1 0.6b · $0.004/free | $0.40 |
Straight arithmetic from published list prices — no volume discounts, batch tiers or committed-use pricing applied. Run your own numbers in the cost calculator.
Provider price floors
Median and floor input price per provider across the models we track, and how many moved this week.
| Provider | Models | Floor in | Median in | Top in | Moved this week |
|---|---|---|---|---|---|
| OpenRouter | 462 | $0.017 | $0.5 | $150 | — |
| Fireworks AI | 273 | $0.05 | $0.22 | $4.5 | — |
| DeepInfra | 135 | $0.019 | $0.27 | $16.5 | — |
| Novita AI | 132 | $0.02 | $0.28 | $4 | — |
| OpenAI | 114 | $0.05 | $2 | $150 | — |
| Azure OpenAI | 111 | $0.05 | $2 | $75 | — |
| Amazon Bedrock | 109 | $0.042 | $0.99 | $18.8 | — |
| Together AI | 100 | $0.02 | $0.525 | $3.5 | — |
| Mistral AI | 76 | $0.06 | $0.4 | $4 | — |
| Databricks | 67 | $0.05 | $1.36 | $30 | — |
| Perplexity | 66 | $0.011 | $1.25 | $10 | — |
| Nebius | 58 | $0.01 | $0.15 | $3 | — |
| Google Gemini | 49 | $0.075 | $0.5 | $2 | — |
| xAI | 44 | $1 | $1.25 | $2 | — |
| Replicate | 40 | $0.03 | $0.65 | $15 | — |
| Cloudflare Workers AI | 30 | $0.017 | $0.351 | $1.92 | — |
| Ollama | 29 | — | — | — | — |
| Anthropic | 28 | $0.25 | $5 | $15 | — |
| Alibaba DashScope | 27 | $0.05 | $0.4 | $2.5 | — |
| Moonshot AI | 24 | $0.2 | $1 | $3 | — |
Frequently asked
› What does the 2026-w39 report measure?
It summarises every price difference detected between consecutive refresh runs of the dataset. Each refresh snapshots input, output and cached-input rates per 1M tokens for every tracked endpoint, then diffs that snapshot against the previous one; the movements listed on this page are those diffs, grouped by ISO week.
› Why does a week sometimes show no changes?
Because providers do not change published rates every week. A report with zero movements means every diff run in that week came back clean — the figures in the market snapshot sections still reflect the current dataset.
› Can I reuse this report?
Yes. The underlying data is published under CC BY 4.0 and served free at /api/models. Attribute Price per 1M with a link when you republish figures.