Rule of thumb: 1 token ≈ ¾ of an English word. A typical chat turn is 1–3K input (with context) and 200–800 output.
| Model | $ / 1M in | $ / 1M out | $ / request | $ / month |
|---|
Notes on this pricing
- Claude Sonnet 5 is introductory pricing through Aug 31, 2026 (standard: $3 / $15 after).
- GPT-5.6 Sol is post-cut promo pricing (through ~Nov 21, 2026); requests over 272K input tokens bill at 2× input / 1.5× output.
- Gemini 3.1 Pro is the ≤200K-context rate; above 200K it rises to $4 / $18. Flash intro pricing doubles Jan 1, 2027.
- Batch APIs are ~50% off at all three vendors; prompt caching cuts repeated input cost up to 90% (Anthropic).
- Open-weights hosted rates (e.g. GLM) vary by provider — treat as ballpark.
Something changed or missing a model? It gets re-verified every Friday for This Week in AIOps.