What AI actually costs
Every number below is the vendor's own published rate, per million tokens, in US dollars. The same job can cost several times less on a different model — that is not a discount we invented, it is the same work run on a different engine.
Model pricing keeps moving — new releases, price cuts, currency swings. Every row is the vendor's own rate as checked on 2026-08-10, not a fixed discount we made up. Before quoting a client, confirm against the vendor's live pricing page.
| Model | Input /1M | Cached /1M | Output /1M | Source |
|---|---|---|---|---|
| Claude Opus 5Anthropic | $5.00 | $0.50 | $25.00 | anthropic ↗ |
| Claude Sonnet 5Anthropic · promo through 2026-08-31, then $3.00 / $0.30 / $15.00 | $2.00 | $0.20 | $10.00 | anthropic ↗ |
| Claude Haiku 4.5Anthropic | $1.00 | $0.10 | $5.00 | anthropic ↗ |
| GPT-5.6 SolOpenAI · flagship tier | $5.00 | $0.50 | $30.00 | openai ↗ |
| GPT-5.6 TerraOpenAI · standard tier | $2.00 | $0.20 | $12.00 | openai ↗ |
| Gemini 3.1 ProGoogle · preview, ≤200k context (above that: $4.00 / $0.40 / $18.00) | $2.00 | $0.20 | $12.00 | google ↗ |
| Model | Input /1M | Cached /1M | Output /1M | Source |
|---|---|---|---|---|
| DeepSeek V4 ProDeepSeek · vendor flags a price increase coming soon | $0.435 | $0.0036 | $0.87 | deepseek ↗ |
| DeepSeek V4 FlashDeepSeek · vendor flags a price increase coming soon | $0.14 | $0.0028 | $0.28 | deepseek ↗ |
| Kimi K3Moonshot · flagship tier, tax not included | $3.00 | $0.30 | $15.00 | kimi ↗ |
| GLM-5.2Zhipu · flagship tier | $1.40 | $0.26 | $4.40 | zhipu ↗ |
| GLM-4.7Zhipu | $0.60 | $0.11 | $2.20 | zhipu ↗ |
| Qwen3-MaxAlibaba Cloud · flagship tier | $2.00 | ≈10% of input | $6.00 | alibaba ↗ |
| Qwen3-PlusAlibaba Cloud · ≤256k context, limited-time 20% off | $0.32 | ≈10% of input | $1.60 | alibaba ↗ |
Not every Chinese model is cheap — Kimi K3 lists at $3.00 / $15.00, level with Claude Sonnet 5's standard rate. The saving comes from picking the right model for the job, not from the vendor's flag.
Longer context can jump to a higher tier — Gemini splits at 200k, Qwen at 32k/256k/1M; the rows above show the lowest tier. Anthropic and OpenAI charge one flat rate across the full context.
Promo prices and flagged increases both have a clock — Claude Sonnet 5's lower rate expires 2026-08-31, and DeepSeek's own page flags an upcoming increase. After the checked date, defer to the vendor's live page.
All figures are each vendor's own published rate, checked 2026-08-10, linked per row. Currency, regional pricing, and enterprise contracts can differ from these — this is the public standard API rate, not what we quote clients.
Already paying for one of these? Tell us your usage and we'll work out what the same job costs on a Chinese model instead.
See the letter we send →