The API pricing log

A dated record of changes in published API prices for frontier language models. A vendor’s pricing page shows today’s rate and nothing else; this page keeps the history. What moved, when, and by how much.

Last updated September 7, 2026. New entries are added as changes are verified.

Entries

How entries are made

Primary sources only. Every entry comes from a vendor’s own pricing page or announcement. An entry is dated by when the change took effect where the vendor names that date, and marked “read” with our verification date otherwise.

Aggregators are not list prices. A reseller’s “cheapest price” reflects routing across providers and can move without the vendor changing anything. Useful signal, but a different quantity; it is never recorded here as a vendor price change.

Slots, not names. Vendors rename and re-ladder models constantly. A price trend has to track the role a model plays (the frontier model, the cheap workhorse), not its label. A new generation is logged as a model evolution, priced against the model it replaces.

Per-token price is not cost per task. Token efficiency and tokenizer changes move real cost independently of the headline rate. Changes of that kind are logged as effective price changes.

Rates are re-verified, not trusted. The DeepSeek entry surfaced in a routine re-check of a rate that had been verified only three weeks earlier; nothing on the vendor’s page announced the change.

Prices are per million tokens; a change reads as before → after. Newest first.

gpt-5.6-sol drops to a promotional rate

Price changeOne model
ModelBeforeAfterChange
gpt-5.6-solstandard rateInput$5$4-20%
Cached input$0.50$0.40-20%
Output$30$20-33%

OpenAI’s page calls the rate promotional, “available at least through November 21, 2026”, and does not name the rate that follows. Batch and flex move with it to $2 / $10; long context to $8 / $30. It is the first in-place cut on a current model in this log, and it lands the same week a model at two and a half times its price is listed above it.

OpenAI’s GPT-6 Astra: the frontier slot doubles on input

Model evolutionFrontier slot
New modelPredecessorNewChange
gpt-6-astra (frontier)listed above gpt-5.6-solInput$5$10+100%
Cached input$0.50$1+100%
Output$30$50+67%

GPT-6 Astra reached the API on September 3 as “our most capable model”, listed above the gpt-5.6 line and recommended as the place to start. Priced against gpt-5.6-sol at its list rate, input doubles, output rises by two thirds, and the cached-input rate doubles. Batch and flex stay at half of standard; past 272k input tokens the rate is $20 / $75. sol stays listed, now at a promotional rate (entry above). Since gpt-5 the frontier slot has gone from $1.25 / $10 to $10 / $50: 8x on input, 5x on output.

Claude Fable 5.1: headline unchanged, cache reads cut by three quarters

Model evolutionFrontier slot
New modelPredecessorNewChange
Claude Fable 5.1succeeds Claude Fable 5Input$10$100%
Cache read$1$0.25-75%
Output$50$500%

Anthropic launched Claude Fable 5.1 on September 1 as the successor to Claude Fable 5, at the same $10 / $50. The change is on the cache: a cache hit costs 2.5% of the input rate, where every other Claude model charges 10%. For a workload that re-sends a long prefix on every call, that is a real cut the headline does not show. At a 90% cache-hit rate the effective input rate falls from $1.90 to $1.23 per million tokens. GPT-6 Astra, at the same list price, keeps a $1 cache read. Fable 5 itself, launched in June at $10 / $50, predates this log; Anthropic’s Opus line stays at $5 / $25.

OpenAI’s gpt-5.6 generation: the current lineup goes from five models to three

Model evolutionFull lineup
New modelPredecessorNewChange
sol (frontier)succeeds gpt-5.5Input$5$50%
Output$30$300%
terra (mid)succeeds gpt-5.4Input$2.50$2-20%
Output$15$12-20%
luna (cheap)succeeds gpt-5.4-nanoInput$0.20$0.200%
Output$1.25$1.20-4%

gpt-5.6 arrives as three models where the prior generation offered five (gpt-5.4, its mini, nano, and pro variants, plus gpt-5.5). Against each slot’s predecessor: the frontier rate holds, the mid slot lists 20% lower on both sides, the cheap slot trims output by 4%. This is a succession, not a removal: as of the read date, the prior generation remained listed at unchanged prices.

Anthropic: a ~30% effective increase at an unchanged headline price

Effective price changeClaude 4.7 and later

Anthropic states that Claude 4.7 and later use a newer tokenizer that produces approximately 30% more tokens for the same text. The per-token price did not change, so the cost per unit of actual work rose about 30% on those models: a price change that appears on no pricing page and that a per-token comparison cannot see.

DeepSeek repriced up roughly 3x and introduced peak and off-peak rates

Price changeFull price list
ModelBeforeAfterChange
deepseek-v4-propeakInput$0.435$1.32+203%
Output$0.87$3.96+355%
off-peakInput$0.435$0.66+52%
Output$0.87$1.98+128%
deepseek-v4-flashpeakInput$0.14$0.44+214%
Output$0.28$1.32+371%
off-peakInput$0.14$0.22+57%
Output$0.28$0.66+136%

Announced and effective 16:00 UTC with the V4-Pro release. Peak hours are 01:00–04:00 and 06:00–10:00 UTC; off-peak is half of peak.

Get new entries by email

One email per entry, nothing else.

Was this helpful?