Verified 2026 moves (first hand, dated)
All prices USD per million tokens. These entries come from our own verification runs and watcher diffs:
- 2026-07-30, OpenAI cuts Terra and Luna.
gpt-5.6-lunafrom $1 / $6 to $0.20 / $1.20 (80 percent), andgpt-5.6-terrafrom $2.50 / $15 to $2 / $12 (20 percent). Luna became the cheapest current OpenAI model on output, under gpt-5.4-nano’s $1.25. Full analysis in the GPT-5.6 price drop. - Verified 2026-07-19, the Opus generation reset.
claude-opus-4-8and 4-7 price at $5 / $25 while the legacyclaude-opus-4line remains $15 / $75: the Opus tier became three times cheaper across generations. Meanwhileclaude-haiku-4-5at $1 / $5 costs more than the 3.5 Haiku generation’s $0.80 / $4, a quiet budget tier increase. Sonnet has held $3 / $15 across three generations. - Verified 2026-07-22, Gemini’s flash line drifted up. The 3.x generation flash models ($1.50 input on
gemini-3.6-flash) sit well above the 2.5 flash era’s $0.30, with the lite tier taking over the old price point. Generation upgrades can be price increases wearing new names. - Verified 2026-07-22, Kimi stays the undercut.
kimi-k2.6at $0.95 / $4 continues Moonshot’s pattern of pricing just below Western mid tiers; the comparison is in Kimi vs Claude vs GPT cost.
The longer arc (approximate, from public announcements)
Reconstructed from provider announcements over the years; treat the exact figures as approximate and the shape as reliable:
- 2023: GPT-4 launches around $30 / $60. GPT-4 Turbo later lands around $10 / $30.
- 2024: GPT-4o steps down to $5 / $15, then $2.50 / $10. GPT-4o-mini arrives near $0.15 / $0.60, replacing the $0.50 / $1.50 GPT-3.5-turbo era. Claude 3 Opus opens at $15 / $75; Claude 3.5 Sonnet at $3 / $15.
- 2025: OpenAI cuts its o3 reasoning model roughly 80 percent mid year, the pattern repeating.
- 2026: the moves in the verified section above, output on a capable flagship line now around $12 versus GPT-4’s $60 three years earlier.
The three lessons in the data
- Deflation is real but tiered. Flagships fell about 5 to 10 times. Budget tiers wobbled and sometimes rose. If you only remember one thing: check your tier, not the headline.
- Changes are steps, not slopes. An 80 percent overnight cut breaks every hardcoded assumption at once: budget alerts, routing rules, margin dashboards. This is why price tables need to be a live feed, not a constant in your code.
- Cuts are opportunities with a deadline. Every step down makes some feature affordable that was not, and re running the math from estimating AI feature cost the week of a cut is how you ship it first.
This page updates when our weekly watcher confirms a change. Last updated 2026-08-07.
FAQ
Do LLM API prices go down over time?
At the flagship level, dramatically: GPT-4 launched around $30 input and $60 output per million tokens in 2023, and current mid flagships sit at $2 to $5 input. But it is not monotonic. Budget tiers have quietly risen in places, Claude Haiku 4.5 at $1 and $5 costs more than the 3.5 Haiku generation did, and Gemini’s flash line is pricier in its 3.x generation than 2.5 was. Direction depends on the tier.
What was the biggest recent LLM price change?
The July 30, 2026 OpenAI cut: GPT-5.6 Luna fell 80 percent to $0.20 input and $1.20 output per million, and GPT-5.6 Terra fell 20 percent to $2 and $12. Luna became the cheapest current OpenAI model on output overnight, undercutting gpt-5.4-nano, and every hardcoded price table referencing the old rates became silently wrong.
Why does price history matter for my product?
Because your unit economics assumptions have a shelf life in both directions. A feature that was unaffordable on a flagship model a year ago may be cheap now, and a budget assumption from last quarter may be stale after one cut. Teams that re run their cost math when prices move ship features their competitors still believe are too expensive.
How can I get notified when a model price changes?
Automate the diff. We run a weekly watcher that compares our price tables against published provider rates and opens a reviewed pull request when something moves, and the resulting table is public as JSON at useweckr.com/pricing.json. Watching that feed, or the LiteLLM community dataset it is diffed against, is the practical notification channel today, since providers do not offer one.
Are historical LLM prices documented anywhere official?
Poorly. Providers update their pricing pages in place, so the old number disappears when the new one lands, and announcements are scattered across blogs and community posts. That is why this page separates what we have verified first hand with dates from the longer arc reconstructed from public announcements, which we label approximate.
Keep reading
History repeats. Your cost tracking should notice.
The next step change is a matter of months. Weckr’s price table is watched weekly, published live, and used to recompute every customer’s real cost, so when the market moves, your margins move with it instead of your spreadsheet lying quietly. See it on the live demo, or start with the AI cost and margin guide.