OpenAI cuts GPT-5.6 API prices by up to 80%, adds a faster paid tier

Luna drops to $0.20 per million input tokens, Terra falls 20%, and a new Sol Fast mode runs 2.5 times quicker for double the price.

ThreatVectr Newsdesk· 4 min read
A glowing digital interface with abstract neural network patterns dissolving into fragmented, distorted code fragments on a dark background, blue and amber ligh
Share

Key points

  • OpenAI cut the API price of GPT-5.6 Luna by 80%, from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output tokens.
  • GPT-5.6 Terra fell 20%, from $2.50 to $2 per million input tokens and from $15 to $12 per million output tokens.
  • Auto-review in the ChatGPT app and Codex command-line tool moves from GPT-5.4 to GPT-5.6 Luna, which OpenAI says cuts that feature's cost by roughly ten times.
  • A new GPT-5.6 Sol Fast mode runs up to 2.5 times faster than standard, at twice the standard API price, aimed at time-sensitive coding and agent workloads.
  • Standard pricing for GPT-5.6 Sol is unchanged.

OpenAI has slashed the prices developers pay to use two of its GPT-5.6 models, and added a paid speed boost for a third. The changes were announced by the company on X and first reported by BleepingComputer.

The biggest cut lands on Luna, the cheaper of the two models. Its price falls by 80%.

Terra, the middle-tier model, gets a 20% cut. Sol, the top model, keeps its standard price but gains a new Fast mode for customers who need answers quicker.

What actually changed in the price list?

OpenAI dropped the per-token price developers pay to send text into these models and receive text back. A token is roughly three-quarters of a word, and prices are quoted per million tokens, which is the industry norm.

Here are the new numbers side by side.

Model Input (per 1M tokens) Output (per 1M tokens) Change
GPT-5.6 Luna $0.20 (was $1) $1.20 (was $6) -80%
GPT-5.6 Terra $2 (was $2.50) $12 (was $15) -20%
GPT-5.6 Sol (standard) unchanged unchanged 0%
GPT-5.6 Sol Fast 2x standard 2x standard +100% for 2.5x speed

The cuts also change how usage is counted inside Codex, OpenAI's coding tool, and ChatGPT Work, the workplace version of ChatGPT. Tasks that use these cheaper models now eat less of a customer's monthly allowance, so the same subscription stretches further.

Why does this matter to people who don't write code?

Cheaper AI calls tend to show up as cheaper or more generous products a few weeks later. If you use a writing assistant, a customer-service chatbot, or a coding tool built on OpenAI's models, the company running it is now paying a fraction of what it did last week.

Some will pass that on. Some will quietly pocket it. Either way, the direction of travel is the same: the raw cost of a machine-generated sentence keeps falling.

One concrete example already in the wild: Auto-review, the feature in ChatGPT and Codex that checks your work, has been switched from the older GPT-5.4 model to GPT-5.6 Luna. OpenAI says that alone cuts the feature's cost by around ten times.

What is Sol Fast mode, and who is it for?

Sol Fast is a new option for developers who call the Sol model through OpenAI's API, the interface software uses to talk to the model. It runs up to 2.5 times faster than the standard version and costs twice as much per token.

OpenAI says the model's intelligence is not reduced in Fast mode. It is aimed at jobs where waiting is expensive: live coding help, research agents that chain many calls together, and any system where a user is watching a spinner.

For most everyday use, standard Sol is fine. Paying double for speed only makes sense when the seconds themselves have a price tag.

Is there a security angle here?

Indirectly, yes. Cheaper models make it cheaper for defenders to run large-scale log review, phishing-email triage and code scanning through an AI. They also make it cheaper for attackers to generate convincing lure emails at volume. The economics cut both ways, and they just got sharper.

© 2026 Threat Vectr