GPT 5.4 vs Gemini 2.0 Flash Lite 001 pricing
Side-by-side token pricing for GPT 5.4 (OpenAI) and Gemini 2.0 Flash Lite 001 (Google Gemini), plus the total cost at three workload sizes so the crossover point is obvious. Prices refreshed 2026-09-13 13:40 UTC.
| Metric | GPT 5.4 | Gemini 2.0 Flash Lite 001 |
|---|---|---|
| Input $/1M tokens | $2.5 | $0.075 |
| Output $/1M tokens | $15 | $0.3 |
| Cached input $/1M | $0.25 | $0.019 |
| Context window | 1.1M | 1.0M |
| Max output | 128K | 8K |
| Vision | yes | yes |
| Tool calling | yes | yes |
| Cache pricing | yes | yes |
Total cost by workload
| Workload | GPT 5.4 | Gemini 2.0 Flash Lite 001 | Cheaper |
|---|---|---|---|
| 1M tokens in · 300K out | $7 | $0.165 | Gemini 2.0 Flash Lite 001 (98% less) |
| 10M tokens in · 3M out | $70 | $1.65 | Gemini 2.0 Flash Lite 001 (98% less) |
| 100M tokens in · 30M out | $700 | $16.5 | Gemini 2.0 Flash Lite 001 (98% less) |
On blended cost (weighted 3:1 toward input), Gemini 2.0 Flash Lite 001 comes out cheaper by roughly 97% on input pricing. Pick Gemini 2.0 Flash Lite 001 when the workload is price-sensitive and its context window covers your prompts; GPT 5.4 is the one to benchmark if you need vision, reasoning depth, or a different quality profile.
Frequently asked
› Is GPT 5.4 cheaper than Gemini 2.0 Flash Lite 001?
On input pricing, Gemini 2.0 Flash Lite 001 is cheaper ($0.075 vs $2.5 per 1M tokens). On blended cost the winner is Gemini 2.0 Flash Lite 001, because output tokens are weighted differently for most workloads.
› Which has the larger context window?
GPT 5.4 offers the larger context (1.1M tokens).
› Do these prices include prompt caching?
Cached input is priced separately where a provider publishes a cache rate; the table above lists it per model so the comparison stays like-for-like.