GPT 5.4 vs Glm 5 2 pricing
Side-by-side token pricing for GPT 5.4 (OpenAI) and Glm 5 2 (Mistral AI), plus the total cost at three workload sizes so the crossover point is obvious. Prices refreshed 2026-09-13 13:40 UTC.
| Metric | GPT 5.4 | Glm 5 2 |
|---|---|---|
| Input $/1M tokens | $2.5 | $1.4 |
| Output $/1M tokens | $15 | $4.4 |
| Cached input $/1M | $0.25 | $0.14 |
| Context window | 1.1M | 1.0M |
| Max output | 128K | 131K |
| Vision | yes | no |
| Tool calling | yes | yes |
| Cache pricing | yes | yes |
Total cost by workload
| Workload | GPT 5.4 | Glm 5 2 | Cheaper |
|---|---|---|---|
| 1M tokens in · 300K out | $7 | $2.72 | Glm 5 2 (61% less) |
| 10M tokens in · 3M out | $70 | $27.2 | Glm 5 2 (61% less) |
| 100M tokens in · 30M out | $700 | $272 | Glm 5 2 (61% less) |
On blended cost (weighted 3:1 toward input), Glm 5 2 comes out cheaper by roughly 44% on input pricing. Pick Glm 5 2 when the workload is price-sensitive and its context window covers your prompts; GPT 5.4 is the one to benchmark if you need vision, reasoning depth, or a different quality profile.
Frequently asked
› Is GPT 5.4 cheaper than Glm 5 2?
On input pricing, Glm 5 2 is cheaper ($1.4 vs $2.5 per 1M tokens). On blended cost the winner is Glm 5 2, because output tokens are weighted differently for most workloads.
› Which has the larger context window?
GPT 5.4 offers the larger context (1.1M tokens).
› Do these prices include prompt caching?
Cached input is priced separately where a provider publishes a cache rate; the table above lists it per model so the comparison stays like-for-like.