Three weeks after launching GPT-5.6, OpenAI slashed the price of its fastest model by 80%. That is not a routine tweak — it is a signal about who now holds the power in AI, and the answer is increasingly “whoever charges less.” On July 30, 2026, OpenAI cut GPT-5.6 Luna from $1 to $0.20 per million input tokens, part of a broader collapse in AI prices driven by fierce competition. Here is exactly what changed in OpenAI API pricing, why it is happening, and what a full-blown AI price war means for you.

What exactly changed

OpenAI’s GPT-5.6 family, launched July 9, 2026, comes in three tiers. The July 30 repricing hit them unevenly:

  • Luna (fast, cost tier): $1/$6 → $0.20/$1.20 per million input/output tokens — an 80% cut on input.
  • Terra (mid tier): $2.50/$15 → $2/$12 — a 20% cut.
  • Sol (flagship): unchanged at $5/$30.

The pattern is telling. OpenAI did not touch the flagship, where it still claims a capability edge and pricing power. It went aggressive precisely on the cheap, high-volume tier — the one that competes head-on with low-cost rivals. That is where the pressure is, and OpenAI knows it.

For context, this came just three weeks after GPT-5.6 launched. Cutting a major model’s price 80% that fast is not something a company does from a position of comfort.

The three GPT-5.6 tiers, explained

If the names Sol, Terra, and Luna are new to you, here is the quick map — it is the key to using OpenAI’s lineup without overpaying:

  • Sol is the flagship: the most capable, most expensive tier at $5/$30 per million tokens. Reach for it on the hardest reasoning, complex coding, and tasks where a wrong answer is costly. It kept its price because that is where OpenAI still claims a clear capability lead.
  • Terra is the balanced middle: now $2/$12 after a 20% cut. It handles most demanding work — solid reasoning, longer analysis, agent steps — at a meaningful discount to the flagship.
  • Luna is the speed-and-cost tier: now just $0.20/$1.20 after the 80% cut. It is built for high volume and everyday tasks — summarizing, drafting, classifying, extracting, simple Q&A — where you want good-enough answers fast and cheap.

The single most useful money habit is to stop defaulting to the biggest model. Most real workloads are Luna and Terra jobs; Sol is for the genuinely hard minority. Getting that split right is often the difference between a large AI bill and a small one.

Why OpenAI blinked: the competition

The repricing is a direct response to competitive pressure, and the numbers behind it are striking.

Reporting around the cut cited data showing Chinese models capturing roughly 46% of US enterprise token usage on OpenRouter, a popular model-routing platform. In other words, nearly half of the tokens flowing through that ecosystem were going to non-US models — many of them cheap and open-weight.

The most visible pressure comes from labs like DeepSeek, whose V4 models pair strong performance with prices a fraction of the closed frontier. When a capable open-weight model costs a fraction of a cent per thousand tokens and can be self-hosted for free, a $1-per-million input price starts to look untenable. By dropping Luna to $0.20 input, OpenAI undercut some rivals on input cost outright — though it can remain pricier on output.

The strategic shift is clear: for a growing slice of the market, the competition is no longer mainly about who has the smartest model. It is about price-performance — the most capability per dollar.

This is a price war, and you are winning it

Zoom out and the Luna cut is one move in a broad, fast repricing of AI:

  • Frontier labs are cutting per-token costs repeatedly, sometimes within weeks of a launch.
  • Open-weight models keep dragging the floor lower, because “free to self-host” is the ultimate price anchor.
  • Even where flagship models hold their price, cheaper tiers are proliferating to capture cost-sensitive users.

Not every move is a cut, which is worth noting so you plan correctly: some intro promotions are also expiring. For example, Claude Sonnet 5’s introductory rate was set to rise from $2 to $3 per million input tokens on September 1, 2026. So the trend is “cheaper on average, but watch the fine print,” not “everything only ever gets cheaper.”

Still, the direction is unambiguous, and it is great news if you use AI: the same work costs a fraction of what it did a year ago, and the cheapest AI model that is genuinely good enough for a task keeps getting better.

It is worth being clear-eyed about the trade behind the trend, though. Aggressive price cuts are partly a race for market share, subsidized by labs betting on volume and lock-in. That is fine for you in the short run — cheaper tokens are cheaper tokens — but it means today’s promotional rate is not a promise. Build your workflows so you can switch providers if a price rises or a better deal appears, rather than hard-wiring your business to one lab’s current pricing.

What it means for you

Whether you are an individual using apps or a business calling APIs, a price war changes the smart playbook.

  • Match the model to the job. You rarely need the $5/$30 flagship. For summarizing, drafting, classifying, and most everyday tasks, a fast tier like Luna at $0.20 input does the work for a fraction of the cost. Reserve the flagship for genuinely hard reasoning.
  • Cost is now a feature you can shop for. With prices moving monthly, it pays to know the current rates and switch when a better price-performance option appears.
  • Cheaper tokens unlock new uses. Things that were too expensive to automate — processing long documents, running agents over many steps, high-volume workflows — become viable as prices fall.
  • Don’t over-optimize on price alone. The cheapest model that fails your task is not cheap. Reliability, speed, context length, and data policies still matter; price is one axis of several.

The users who benefit most are the ones who understand the lineup — which model, which tier, for which task — instead of defaulting to the most expensive option out of habit.

A simple way to cut your AI costs today

You do not need a procurement analysis to benefit from the price war. A few practical moves capture most of the savings:

  • Audit your default. If you (or your tools) send everything to the most expensive model, you are almost certainly overpaying. Route routine tasks to a cheap, fast tier.
  • Tier by difficulty, not by habit. Draft, summarize, and classify on the cheap tier; escalate to a mid or flagship model only when a task genuinely needs deeper reasoning.
  • Use caching where you can. Many providers charge far less for repeated or cached context, which adds up fast in agents and long workflows.
  • Re-check prices periodically. Rates are moving monthly in 2026; a model that was expensive last quarter may be a bargain now, and vice versa as intro offers expire.
  • Measure quality, not just cost. Track whether the cheaper model actually does the job. The goal is the best result per dollar, not the lowest sticker price.

Done consistently, matching the model to the task is the highest-leverage cost habit in AI right now — and the price war only makes it more valuable.

Turn cheaper AI into real leverage — with Coursiv

Falling prices only help if you know how to use the tools they unlock. The difference between reading about an 80% price cut and actually building cheaper, faster workflows on top of it is skill — and that is what Coursiv delivers: practical, plain-English AI training that teaches you to get results from the major tools without wasting money or effort.

Coursiv’s guided lessons help you understand tokens, models, and pricing tiers, then apply the right tool to real tasks in your work or business. If the AI price war made you realize how much capability is now within reach, the next step is learning to wield it deliberately: start upskilling with Coursiv today.

Final verdict

OpenAI cutting GPT-5.6 Luna by 80% just three weeks after launch is the clearest sign yet that AI’s competitive battleground has shifted from pure capability to price-performance — driven by cheap, open-weight rivals eating into enterprise usage. For users, that is close to unambiguously good: frontier-grade AI keeps getting cheaper, and matching the right tier to each task now saves real money. Just keep an eye on the fine print, since some intro rates expire even as others fall. The winning habit in a price war is fluency — knowing the lineup well enough to always pick the most capability per dollar.

FAQ

How much did OpenAI cut GPT-5.6 prices?
On July 30, 2026, OpenAI cut Luna 80% on input, from $1/$6 to $0.20/$1.20 per million input/output tokens, and Terra 20%, from $2.50/$15 to $2/$12. The flagship Sol stayed at $5/$30.
Why did OpenAI lower its prices?
Competition, especially from cheap Chinese open-weight models like DeepSeek that captured a large share of enterprise token usage. OpenAI cut its cost tier hardest to compete on price-performance, where the pressure is greatest.
What is the cheapest GPT-5.6 model now?
Luna is the cheapest tier at $0.20 per million input and $1.20 per million output tokens. It is aimed at fast, high-volume, cost-sensitive tasks rather than the hardest reasoning, which the flagship Sol tier targets.
Is AI always getting cheaper?
On average, yes — competition and open-weight models keep pushing prices down. But watch the fine print: some introductory rates expire, such as Claude Sonnet 5’s planned increase from $2 to $3 per million input tokens on September 1, 2026.
How do I save money using AI?
Match the model to the task: use cheaper, fast tiers for everyday work and reserve expensive flagships for genuinely hard problems. Track current pricing, since it changes often, and weigh reliability and speed alongside cost.
What are Sol, Terra, and Luna?
They are the three tiers of OpenAI’s GPT-5.6 family. Sol is the top-capability, top-price flagship ($5/$30 per million tokens); Terra is the balanced mid tier ($2/$12 after a cut); and Luna is the fast, cheap tier ($0.20/$1.20 after an 80% input cut).
Should I switch to a cheaper AI model?
Often, yes — for routine tasks. Most everyday work does not need a flagship model, so routing it to a cheaper, fast tier saves money with little quality loss. Reserve premium models for genuinely hard reasoning, and always confirm the cheaper option actually does the job.