Every AI model price, per 1M tokens.
1,818 priced chat models and 1,907 endpoints total, from 33 providers. Compare cost per 1M tokens, prompt-cache rates and context windows — then estimate your own monthly bill in the cost calculator.
Search the index
Filter by provider, sort by what your workload actually pays for.
| Model | Provider | Input $/1M | Output $/1M | Context |
|---|---|---|---|---|
| Gemini Exp 1114 | Google Gemini | free | free | 1.0M |
| Gemini Exp 1206 | Google Gemini | free | free | 2.1M |
| Gemma 3 27b It | Google Gemini | free | free | 131K |
| Gemma 4 26b A4b It | Google Gemini | free | free | 262K |
| Gemma 4 31b It | Google Gemini | free | free | 262K |
| Learnlm 1.5 Pro Experimental | Google Gemini | free | free | 33K |
| Lyria 3 Clip Preview | Google Gemini | free | free | 131K |
| Lyria 3 Pro Preview | Google Gemini | free | free | 131K |
| Lyria 3.5 Clip Preview | Google Gemini | free | free | 131K |
| Lyria 3.5 Pro Preview | Google Gemini | free | free | 131K |
| Lyria 3.5 | Google Gemini | free | free | 1.0M |
| Labs Leanstral 1 5 | Mistral AI | free | free | 262K |
| Labs Leanstral 1 5 1 | Mistral AI | free | free | 262K |
| Pplx 70b Online | Perplexity | free | $2.8 | 4K |
| Pplx 7b Online | Perplexity | free | $0.28 | 4K |
| Sonar Medium Online | Perplexity | free | $1.8 | 12K |
| Sonar Small Online | Perplexity | free | $0.28 | 12K |
| Anthropic.claude Mythos Preview | Amazon Bedrock | free | free | 1M |
| Gemma 2b It Lora | Cloudflare Workers AI | free | free | 8K |
| Mistral 7b Instruct V0.2 Lora | Cloudflare Workers AI | free | free | 15K |
| Llama 2 7b Chat Hf Lora | Cloudflare Workers AI | free | free | 8K |
| Gemma 7b It Lora | Cloudflare Workers AI | free | free | 4K |
| Codegeex4 | Ollama | free | free | 33K |
| Codegemma | Ollama | free | free | 8K |
| Codellama | Ollama | free | free | 4K |
| Deepseek Coder V2 Base | Ollama | free | free | 8K |
| Deepseek Coder V2 Instruct | Ollama | free | free | 33K |
| Deepseek Coder V2 Lite Base | Ollama | free | free | 8K |
| Deepseek Coder V2 Lite Instruct | Ollama | free | free | 33K |
| Deepseek V3.1:671b Cloud | Ollama | free | free | 164K |
| GPT Oss:120b Cloud | Ollama | free | free | 131K |
| GPT Oss:20b Cloud | Ollama | free | free | 131K |
| Internlm2 5 20b Chat | Ollama | free | free | 33K |
| Llama2 | Ollama | free | free | 4K |
| Llama2 Uncensored | Ollama | free | free | 4K |
| Llama2:13b | Ollama | free | free | 4K |
| Llama2:70b | Ollama | free | free | 4K |
| Llama2:7b | Ollama | free | free | 4K |
| Llama3 | Ollama | free | free | 8K |
| Llama3.1 | Ollama | free | free | 8K |
| Llama3:70b | Ollama | free | free | 8K |
| Llama3:8b | Ollama | free | free | 8K |
| Mistral | Ollama | free | free | 8K |
| Mistral 7B Instruct V0.1 | Ollama | free | free | 8K |
| Mistral 7B Instruct V0.2 | Ollama | free | free | 33K |
| Mistral Large Instruct 2407 | Ollama | free | free | 66K |
| Mixtral 8x22B Instruct V0.1 | Ollama | free | free | 66K |
| Mixtral 8x7B Instruct V0.1 | Ollama | free | free | 33K |
| Orca Mini | Ollama | free | free | 4K |
| Qwen3 Coder:480b Cloud | Ollama | free | free | 262K |
| Vicuna | Ollama | free | free | 2K |
| Auto | OpenRouter | free | free | 2M |
| Free | OpenRouter | free | free | 200K |
| Bodybuilder | OpenRouter | free | free | 128K |
| Nemotron 3.5 Lightning:free | OpenRouter | free | free | 1M |
| Laguna S 2.1:free | OpenRouter | free | free | 262K |
| Laguna Xs 2.1:free | OpenRouter | free | free | 262K |
| Glm 5.2:free | OpenRouter | free | free | 256K |
| Nemotron 3.5 Content Safety:free | OpenRouter | free | free | 128K |
| Nemotron 3 Ultra 550b A55b:free | OpenRouter | free | free | 1M |
| Minimax M3:free | OpenRouter | free | free | 1.0M |
| Nemotron 3 Nano Omni 30b A3b Reasoning:free | OpenRouter | free | free | 256K |
| Gemma 4 26b A4b It:free | OpenRouter | free | free | 262K |
| Gemma 4 31b It:free | OpenRouter | free | free | 262K |
| Minimax M2.7:free | OpenRouter | free | free | 197K |
| Nemotron 3 Super 120b A12b:free | OpenRouter | free | free | 262K |
| Llama 3.3 70B Instruct Turbo Free | Together AI | free | free | — |
| Ternary Bonsai 27B | Together AI | free | free | 262K |
| Flux 1 Dev Controlnet Union | Fireworks AI | $0.001 | $0.001 | 4K |
| Qwen2.5 Coder 7B | Nebius | $0.01 | $0.03 | 33K |
| Qwen2.5 Coder 3B Instruct | Nscale | $0.01 | $0.03 | — |
| Qwen2.5 Coder 7B Instruct | Nscale | $0.01 | $0.03 | — |
| Nemotron 3.5 Lightning 30b A3b | Perplexity | $0.011 | $0.17 | — |
| Granite 4.0 H Micro | Cloudflare Workers AI | $0.017 | $0.112 | 131K |
| Mistral Nemo Instruct 2407 | DeepInfra | $0.019 | $0.03 | 131K |
| Mistral Nemo | OpenRouter | $0.019 | $0.03 | 131K |
| Llama 3.2 3B Instruct | DeepInfra | $0.02 | $0.02 | 131K |
| Meta Llama 3.1 8B Instruct Turbo | DeepInfra | $0.02 | $0.04 | 131K |
| Gemma 4 E4B It | DeepInfra | $0.02 | $0.1 | 131K |
| Llama Guard 3 8B | Nebius | $0.02 | $0.06 | 128K |
| Meta Llama 3.1 8B Instruct | Nebius | $0.02 | $0.06 | 128K |
| Qwen2 VL 7B Instruct | Nebius | $0.02 | $0.06 | 131K |
| Paddleocr Vl | Novita AI | $0.02 | $0.02 | 16K |
| Llama 3.1 8b Instruct | Novita AI | $0.02 | $0.05 | 16K |
| Llama 3.2 1b Instruct | Novita AI | $0.02 | $0.02 | 131K |
| DeepSeek R1 Distill Llama 8B | Nscale | $0.025 | $0.025 | — |
| Llama 3.2 1b Instruct | Cloudflare Workers AI | $0.027 | $0.201 | 60K |
| Llama 3.2 1b Instruct | OpenRouter | $0.027 | $0.201 | 60K |
| Meta Llama 3 8B Instruct | DeepInfra | $0.03 | $0.06 | 8K |
| Meta Llama 3.1 8B Instruct | DeepInfra | $0.03 | $0.05 | 131K |
| GPT Oss 20b | DeepInfra | $0.03 | $0.14 | 131K |
| Llama Prompt Guard 2 22m | Groq | $0.03 | $0.03 | 512 |
| Deepseek Ocr | Novita AI | $0.03 | $0.03 | 8K |
| Qwen3 4b Fp8 | Novita AI | $0.03 | $0.03 | 128K |
| Llama 3.2 3b Instruct | Novita AI | $0.03 | $0.05 | 33K |
| Deepseek Ocr 2 | Novita AI | $0.03 | $0.03 | 8K |
| Llama 3.1 8B Instruct | Nscale | $0.03 | $0.03 | — |
| GPT Oss 20b | OpenRouter | $0.03 | $0.13 | 131K |
| Qwen3.7 Flash | OpenRouter | $0.03 | $0.13 | 1M |
| Granite 3.3 8b Instruct | Replicate | $0.03 | $0.25 | — |
| Autoglm Phone 9b Multilingual | Novita AI | $0.035 | $0.138 | 66K |
| Qwen3 8b Fp8 | Novita AI | $0.035 | $0.138 | 128K |
| GPT Oss 120b | DeepInfra | $0.037 | $0.17 | 131K |
| GPT Oss 120b | OpenRouter | $0.037 | $0.17 | 131K |
| Qwen2.5 7B Instruct | DeepInfra | $0.04 | $0.1 | 33K |
| L3 8B Lunaris V1 Turbo | DeepInfra | $0.04 | $0.05 | 8K |
| NVIDIA Nemotron Nano 9B V2 | DeepInfra | $0.04 | $0.16 | 131K |
| Llama Prompt Guard 2 86m | Groq | $0.04 | $0.04 | 512 |
| Mistral Nemo Instruct 2407 | Nebius | $0.04 | $0.12 | 128K |
| GPT Oss 20b | Novita AI | $0.04 | $0.15 | 131K |
| Mistral Nemo | Novita AI | $0.04 | $0.17 | 60K |
| Llama 3 8b Instruct | Novita AI | $0.04 | $0.04 | 8K |
| Meta Llama 3.2 1B Instruct | SambaNova | $0.04 | $0.08 | 16K |
| Amazon.nova Micro V1:0 | Amazon Bedrock | $0.042 | $0.168 | 128K |
| Gemma 4 26b A4b It | OpenRouter | $0.042 | $0.22 | 262K |
| Llama 3.2 11b Vision Instruct | Cloudflare Workers AI | $0.049 | $0.676 | 128K |
| Llama 3.2 11B Vision Instruct | DeepInfra | $0.049 | $0.049 | 131K |
| Databricks GPT 5 Nano | Databricks | $0.05 | $0.4 | 272K |
| Qwen Turbo | Alibaba DashScope | $0.05 | $0.2 | 129K |
| Qwen Turbo | Alibaba DashScope | $0.05 | $0.2 | 1M |
| Qwen Turbo | Alibaba DashScope | $0.05 | $0.2 | 1M |
| Qwen Turbo | Alibaba DashScope | $0.05 | $0.2 | 1M |
| GPT 5 Nano | OpenAI | $0.05 | $0.4 | 272K |
| GPT 5 Nano | OpenAI | $0.05 | $0.4 | 272K |
| GPT 5 Nano | Azure OpenAI | $0.05 | $0.4 | 272K |
| Gemma 3 12b It | DeepInfra | $0.05 | $0.15 | 131K |
| Gemma 3 4b It | DeepInfra | $0.05 | $0.1 | 131K |
| Mistral Small 24B Instruct 2501 | DeepInfra | $0.05 | $0.08 | 33K |
| Nemotron 3 Nano 30B A3B | DeepInfra | $0.05 | $0.2 | 262K |
| Nemotron Lightning 3p5 30b A3b | Fireworks AI | $0.05 | $0.2 | 262K |
| Llama 3.1 8b Instant | Groq | $0.05 | $0.08 | 131K |
| Gemma 7b It | Groq | $0.05 | $0.08 | 8K |
| GPT Oss 120b | Novita AI | $0.05 | $0.25 | 131K |
| Gemma 3 12b It | Novita AI | $0.05 | $0.1 | 131K |
| L3 8b Lunaris | Novita AI | $0.05 | $0.05 | 8K |
| L3 8B Stheno V3.2 | Novita AI | $0.05 | $0.05 | 8K |
| Nemotron 3 Nano 30b A3b | Novita AI | $0.05 | $0.2 | 262K |
| GPT 5 Nano | OpenRouter | $0.05 | $0.4 | 272K |
| Nemotron 3 Nano 30b A3b | OpenRouter | $0.05 | $0.2 | 262K |
| Gemma 3 4b It | OpenRouter | $0.05 | $0.1 | 131K |
| Gemma 3 12b It | OpenRouter | $0.05 | $0.15 | 131K |
| Mistral Small 24b Instruct 2501 | OpenRouter | $0.05 | $0.08 | 33K |
| Llama 3.2 3b Instruct | OpenRouter | $0.05 | $0.33 | 131K |
| Llama 3.1 8b Instruct | OpenRouter | $0.05 | $0.08 | 131K |
| Llama 2 7b | Replicate | $0.05 | $0.25 | 4K |
| Llama 2 7b Chat | Replicate | $0.05 | $0.25 | 4K |
| Llama 3 8b | Replicate | $0.05 | $0.25 | 8K |
| Llama 3 8b Instruct | Replicate | $0.05 | $0.25 | 8K |
| Mistral 7b Instruct V0.2 | Replicate | $0.05 | $0.25 | 4K |
| Mistral 7b V0.1 | Replicate | $0.05 | $0.25 | 4K |
| GPT 5 Nano | Replicate | $0.05 | $0.4 | — |
| GPT Oss 20b | Together AI | $0.05 | $0.2 | 131K |
| Llama 3.2 3b Instruct | Cloudflare Workers AI | $0.051 | $0.335 | 80K |
| Qwen3 30b A3b Fp8 | Cloudflare Workers AI | $0.051 | $0.335 | 33K |
| GPT 5 Nano | Azure OpenAI | $0.055 | $0.44 | 272K |
| Llama Guard 3 8B | DeepInfra | $0.055 | $0.055 | 131K |
| Mistral Small 3 2 2506 | Mistral AI | $0.06 | $0.18 | 131K |
| GLM 4.7 Flash | DeepInfra | $0.06 | $0.4 | 203K |
| Ling 3.0 Flash | DeepInfra | $0.06 | $0.18 | 131K |
| Qwen2.5 32B Instruct | Nebius | $0.06 | $0.2 | 128K |
| NVIDIA Nemotron 3 Nano 30B A3B | Nebius | $0.06 | $0.24 | 262K |
| Nemotron 3 Nano Omni | Nebius | $0.06 | $0.24 | 262K |
| Nemotron 3 5 Lightning | Nebius | $0.06 | $0.24 | 1.0M |
| Deepseek R1 0528 Qwen3 8b | Novita AI | $0.06 | $0.09 | 128K |
| Ling 3.0 Flash Fast | Novita AI | $0.06 | $0.18 | 262K |
| Ling 3.0 Flash | Novita AI | $0.06 | $0.18 | 262K |
| Qwen2.5 Coder 32B Instruct | Nscale | $0.06 | $0.2 | — |
| Mythomax L2 13b | OpenRouter | $0.06 | $0.06 | 8K |
| Glm 4.7 Flash | OpenRouter | $0.06 | $0.4 | 200K |
| Laguna Xs 2.1 | OpenRouter | $0.06 | $0.12 | 262K |
| Gemma 3n E4B It | Together AI | $0.06 | $0.12 | 33K |
| NVIDIA Nemotron Nano 9B V2 | Together AI | $0.06 | $0.25 | 131K |
| Glm 4.7 Flash | Cloudflare Workers AI | $0.06 | $0.4 | 131K |
| Qwen3.5 Flash 02 23 | OpenRouter | $0.065 | $0.26 | 1M |
| Deepseek V4 Flash 0731 | OpenRouter | $0.065 | $0.18 | 1.3M |
| Mistral 7b Instruct | Perplexity | $0.07 | $0.28 | 4K |
| Mixtral 8x7b Instruct | Perplexity | $0.07 | $0.28 | 4K |
| Pplx 7b Chat | Perplexity | $0.07 | $0.28 | 8K |
| Sonar Small Chat | Perplexity | $0.07 | $0.28 | 16K |
| Databricks GPT Oss 20b | Databricks | $0.07 | $0.3 | 131K |
| Phi 4 | DeepInfra | $0.07 | $0.14 | 16K |
| Gemma 4 26B A4B It | DeepInfra | $0.07 | $0.34 | 262K |
| GPT Oss 20b | Fireworks AI | $0.07 | $0.3 | 131K |
| Qwen3 Coder 30b A3b Instruct | Novita AI | $0.07 | $0.27 | 160K |
| Ernie 4.5 21B A3b Thinking | Novita AI | $0.07 | $0.28 | 131K |
| Baichuan M2 32b | Novita AI | $0.07 | $0.07 | 131K |
| Ernie 4.5 21B A3b | Novita AI | $0.07 | $0.28 | 120K |
| Qwen2.5 7b Instruct | Novita AI | $0.07 | $0.07 | 32K |
| Glm 4.7 Flash | Novita AI | $0.07 | $0.4 | 200K |
| DeepSeek R1 Distill Qwen 14B | Nscale | $0.07 | $0.07 | — |
| Qwen3 Coder 30b A3b Instruct | OpenRouter | $0.07 | $0.28 | 262K |
| Amazon.nova Lite V1:0 | Amazon Bedrock | $0.072 | $0.288 | 300K |
| Nvidia.nemotron Nano 3 30b | Amazon Bedrock | $0.072 | $0.288 | 262K |
| Nvidia.nemotron Nano 9b V2 | Amazon Bedrock | $0.072 | $0.276 | 128K |
| Gemini 2.0 Flash Lite | Google Gemini | $0.075 | $0.3 | 1.0M |
| Gemini 2.0 Flash Lite 001 | Google Gemini | $0.075 | $0.3 | 1.0M |
| Mistral Small 3.2 24B Instruct 2506 | DeepInfra | $0.075 | $0.2 | 128K |
| GPT Oss 20b | Groq | $0.075 | $0.3 | 131K |
| GPT Oss Safeguard 20b | Groq | $0.075 | $0.3 | 131K |
| Mistral Small 3.2 24b Instruct | OpenRouter | $0.075 | $0.2 | 128K |
| GPT Oss Safeguard 20b | OpenRouter | $0.075 | $0.3 | 131K |
| Qwen3 32B | DeepInfra | $0.08 | $0.28 | 41K |
| Gemma 3 27b It | DeepInfra | $0.08 | $0.16 | 131K |
| NVIDIA Nemotron 3.5 Lightning | DeepInfra | $0.08 | $0.2 | 262K |
| DeepSeek V4 Flash 0731 | DeepInfra | $0.08 | $0.18 | 1.0M |
| Qwen3 14B | Nebius | $0.08 | $0.24 | 33K |
| Qwen3 4B | Nebius | $0.08 | $0.24 | 33K |
| Qwen3 Vl 8b Instruct | Novita AI | $0.08 | $0.5 | 131K |
| Nemotron 3.5 Lightning | OpenRouter | $0.08 | $0.2 | 262K |
| Qwen3 32b | OpenRouter | $0.08 | $0.28 | 131K |
| Gemma 3 27b It | OpenRouter | $0.08 | $0.45 | 131K |
| Meta Llama 3.2 3B Instruct | SambaNova | $0.08 | $0.16 | 4K |
| Openai.gpt Oss 20b 1:0 | Amazon Bedrock | $0.084 | $0.36 | 128K |
| NVIDIA Nemotron 3 Super 120B A12B | DeepInfra | $0.085 | $0.4 | 262K |
| Nemotron 3 Super 120b A12b | OpenRouter | $0.085 | $0.4 | 1M |
| Deepseek V4 Flash | OpenRouter | $0.085 | $0.171 | 1.0M |
| Qwen3 235B A22B Instruct 2507 | DeepInfra | $0.09 | $0.55 | 262K |
| Qwen3 Next 80B A3B Instruct | DeepInfra | $0.09 | $1.1 | 262K |
| Gemma 4 31B It Turbo | DeepInfra | $0.09 | $0.34 | 262K |
| DeepSeek V4 Flash | DeepInfra | $0.09 | $0.18 | 1.0M |
| Qwen3 235b A22b Instruct 2507 | Novita AI | $0.09 | $0.58 | 131K |
| Qwen3 30b A3b Fp8 | Novita AI | $0.09 | $0.45 | 41K |
| Mythomax L2 13b | Novita AI | $0.09 | $0.09 | 4K |
| DeepSeek R1 Distill Qwen 1.5B | Nscale | $0.09 | $0.09 | — |
| Llama 4 Scout 17B 16E Instruct | Nscale | $0.09 | $0.29 | — |
| Laguna S 2.1 | OpenRouter | $0.09 | $0.18 | 1.0M |
| Gemma 4 31b It | OpenRouter | $0.09 | $0.34 | 262K |
| Qwen3 Next 80b A3b Instruct | OpenRouter | $0.09 | $1.1 | 262K |
| Qwen3 30b A3b Instruct 2507 | OpenRouter | $0.09 | $0.3 | 262K |
| GPT Oss 20b | Replicate | $0.09 | $0.36 | — |
| Gemini 2.0 Flash | Google Gemini | $0.1 | $0.4 | 1.0M |
| Gemini 2.0 Flash 001 | Google Gemini | $0.1 | $0.4 | 1.0M |
| Gemini 2.5 Flash Lite | Google Gemini | $0.1 | $0.4 | 1.0M |
| Gemini 2.5 Flash Lite Preview 09 2025 | Google Gemini | $0.1 | $0.4 | 1.0M |
| Gemini Flash Lite | Google Gemini | $0.1 | $0.4 | 1.0M |
| Gemini 2.5 Flash Lite Preview 06 17 | Google Gemini | $0.1 | $0.4 | 1.0M |
| Muse Spark 1.2 Contributor | Meta | $0.1 | $0.2 | 1.0M |
| Muse Spark 1.3 Contributor | Meta | $0.1 | $0.2 | 1.0M |
| Devstral Small 2505 | Mistral AI | $0.1 | $0.3 | 128K |
| Devstral Small 2507 | Mistral AI | $0.1 | $0.3 | 128K |
| Devstral Small | Mistral AI | $0.1 | $0.3 | 256K |
| Labs Devstral Small 2512 | Mistral AI | $0.1 | $0.3 | 256K |
| Ministral 3b 2512 | Mistral AI | $0.1 | $0.1 | 131K |
| Ministral 3b | Mistral AI | $0.1 | $0.1 | 131K |
| Voxtral Small 2507 | Mistral AI | $0.1 | $0.4 | 33K |
| Voxtral Small | Mistral AI | $0.1 | $0.4 | 33K |
| Mistral Small | Mistral AI | $0.1 | $0.3 | 32K |
| Ministral 3 3b 2512 | Mistral AI | $0.1 | $0.1 | 131K |
| GPT 4.1 Nano | OpenAI | $0.1 | $0.4 | 1.0M |
| GPT 4.1 Nano | OpenAI | $0.1 | $0.4 | 1.0M |
| GPT 4.1 Nano | Azure OpenAI | $0.1 | $0.4 | 1.0M |
| GPT 4.1 Nano | Azure OpenAI | $0.1 | $0.4 | 1.0M |
| GPT Oss 120b | Baseten | $0.1 | $0.5 | — |
| Meta.llama3 2 1b Instruct V1:0 | Amazon Bedrock | $0.1 | $0.1 | 128K |
| Us.meta.llama3 2 1b Instruct V1:0 | Amazon Bedrock | $0.1 | $0.1 | 128K |
| Llama3.1 8b | Cerebras | $0.1 | $0.1 | 128K |
| Gemma 4 26b A4b It | Cloudflare Workers AI | $0.1 | $0.3 | 256K |
| Gemini 2.0 Flash 001 | DeepInfra | $0.1 | $0.4 | 1M |
| Llama 3.3 70B Instruct Turbo | DeepInfra | $0.1 | $0.32 | 131K |
| Llama 4 Scout 17B 16E Instruct | DeepInfra | $0.1 | $0.3 | 328K |
| Llama 3.3 Nemotron Super 49B V1.5 | DeepInfra | $0.1 | $0.4 | 131K |
| Qwen3.6 35B A3B | DeepInfra | $0.1 | $0.95 | 262K |
| Seed 2.0 Mini | DeepInfra | $0.1 | $0.4 | 256K |
| Qwen3.5 9B | DeepInfra | $0.1 | $0.15 | 262K |
| Llama V3p1 8b Instruct | Fireworks AI | $0.1 | $0.1 | 16K |
| Llama V3p2 1b Instruct | Fireworks AI | $0.1 | $0.1 | 16K |
| Llama V3p2 3b Instruct | Fireworks AI | $0.1 | $0.1 | 16K |
| Codegemma 2b | Fireworks AI | $0.1 | $0.1 | 8K |
| Cogito V1 Preview Llama 3b | Fireworks AI | $0.1 | $0.1 | 131K |
| Deepseek Coder 1b Base | Fireworks AI | $0.1 | $0.1 | 16K |
| Deepseek R1 Distill Qwen 1p5b | Fireworks AI | $0.1 | $0.1 | 131K |
| Ernie 4p5 21b A3b Pt | Fireworks AI | $0.1 | $0.1 | 4K |
| Ernie 4p5 300b A47b Pt | Fireworks AI | $0.1 | $0.1 | 4K |
| Flux 1 Dev | Fireworks AI | $0.1 | $0.1 | 4K |
| Flux 1 Schnell | Fireworks AI | $0.1 | $0.1 | 4K |
| Gemma 2b It | Fireworks AI | $0.1 | $0.1 | 8K |
| Llama Guard 3 1b | Fireworks AI | $0.1 | $0.1 | 131K |
| Llama V2 70b | Fireworks AI | $0.1 | $0.1 | 4K |
| Llama V3p1 405b Instruct Long | Fireworks AI | $0.1 | $0.1 | 4K |
| Llama V3p1 70b Instruct 1b | Fireworks AI | $0.1 | $0.1 | 4K |
| Llama V3p2 1b | Fireworks AI | $0.1 | $0.1 | 131K |
| Llama V3p2 3b | Fireworks AI | $0.1 | $0.1 | 131K |
| Minimax M1 80k | Fireworks AI | $0.1 | $0.1 | 4K |
| Ministral 3 3b Instruct 2512 | Fireworks AI | $0.1 | $0.1 | 256K |
| Nemotron Nano V2 12b Vl | Fireworks AI | $0.1 | $0.1 | 4K |
| Phi 2 3b | Fireworks AI | $0.1 | $0.1 | 2K |
| Phi 3 Mini 128k Instruct | Fireworks AI | $0.1 | $0.1 | 131K |
| Qwen2 Vl 2b Instruct | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 0p5b Instruct | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 1p5b Instruct | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 Coder 0p5b | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 Coder 0p5b Instruct | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 Coder 1p5b | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 Coder 1p5b Instruct | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 Coder 3b | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen2p5 Coder 3b Instruct | Fireworks AI | $0.1 | $0.1 | 33K |
| Qwen3 0p6b | Fireworks AI | $0.1 | $0.1 | 41K |
| Qwen3 1p7b | Fireworks AI | $0.1 | $0.1 | 131K |
| Qwen3 1p7b Fp8 Draft | Fireworks AI | $0.1 | $0.1 | 262K |
| Qwen3 1p7b Fp8 Draft 131072 | Fireworks AI | $0.1 | $0.1 | 131K |
Showing 300 of 1,818 matching models (25+ interactive results).
Cheapest input tokens
What the input side of your bill costs per 1M tokens.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini Exp 1114 | Google Gemini | free | free | — | 1.0M |
| 2 | Gemini Exp 1206 | Google Gemini | free | free | — | 2.1M |
| 3 | Gemma 3 27b It | Google Gemini | free | free | — | 131K |
| 4 | Gemma 4 26b A4b It | Google Gemini | free | free | — | 262K |
| 5 | Gemma 4 31b It | Google Gemini | free | free | — | 262K |
| 6 | Learnlm 1.5 Pro Experimental | Google Gemini | free | free | — | 33K |
| 7 | Lyria 3 Clip Preview | Google Gemini | free | free | — | 131K |
| 8 | Lyria 3 Pro Preview | Google Gemini | free | free | — | 131K |
| 9 | Lyria 3.5 Clip Preview | Google Gemini | free | free | — | 131K |
| 10 | Lyria 3.5 Pro Preview | Google Gemini | free | free | — | 131K |
| 11 | Lyria 3.5 | Google Gemini | free | free | — | 1.0M |
| 12 | Labs Leanstral 1 5 | Mistral AI | free | free | — | 262K |
| 13 | Labs Leanstral 1 5 1 | Mistral AI | free | free | — | 262K |
| 14 | Pplx 70b Online | Perplexity | free | $2.8 | — | 4K |
| 15 | Pplx 7b Online | Perplexity | free | $0.28 | — | 4K |
Best blended cost
Weighted 3:1 in favour of input, which is how most apps actually bill.
| # | Model | Provider | Input $/1M | Output $/1M | Cached $/1M | Context |
|---|---|---|---|---|---|---|
| 1 | Gemini Exp 1114 | Google Gemini | free | free | — | 1.0M |
| 2 | Gemini Exp 1206 | Google Gemini | free | free | — | 2.1M |
| 3 | Gemma 3 27b It | Google Gemini | free | free | — | 131K |
| 4 | Gemma 4 26b A4b It | Google Gemini | free | free | — | 262K |
| 5 | Gemma 4 31b It | Google Gemini | free | free | — | 262K |
| 6 | Learnlm 1.5 Pro Experimental | Google Gemini | free | free | — | 33K |
| 7 | Lyria 3 Clip Preview | Google Gemini | free | free | — | 131K |
| 8 | Lyria 3 Pro Preview | Google Gemini | free | free | — | 131K |
| 9 | Lyria 3.5 Clip Preview | Google Gemini | free | free | — | 131K |
| 10 | Lyria 3.5 Pro Preview | Google Gemini | free | free | — | 131K |
| 11 | Lyria 3.5 | Google Gemini | free | free | — | 1.0M |
| 12 | Labs Leanstral 1 5 | Mistral AI | free | free | — | 262K |
| 13 | Labs Leanstral 1 5 1 | Mistral AI | free | free | — | 262K |
| 14 | Anthropic.claude Mythos Preview | Amazon Bedrock | free | free | — | 1M |
| 15 | Gemma 2b It Lora | Cloudflare Workers AI | free | free | — | 8K |
Cost by use case
Pre-ranked shortlists — each page is computed from the raw dataset, not written by hand.
Providers
Model counts and price floors per provider.
Frequently asked
› How often is this pricing data updated?
The dataset is refreshed on an automated schedule (every 6 hours by default) from public provider data and community-maintained price registries. Each model page and the JSON API carry the timestamp of the last successful refresh.
› Where do these numbers come from?
Every price keeps its provenance: each model page lists the upstream sources it was assembled from, and where two independent sources overlap the page shows both so you can see whether they agree. See the methodology page for the full pipeline.
› Is this the price I will actually pay?
No. Providers change prices, add regional multipliers, tiered discounts and minimum commitments. Use these figures to compare models on a like-for-like basis, then confirm with the provider's own pricing page before you commit to a budget.
› Can I use this data in my own project?
Yes — the JSON API at /api/models is free to read, and llms.txt plus llms-full.txt expose the same dataset for AI agents. Attribution is appreciated.