On 22 September 2026 OpenAI filled in the lower two tiers of its GPT-6 family. GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens. GPT-6 Luna costs $0.10 and $0.50. Both replace their GPT-5.6 namesakes at roughly half the price, and OpenAI has described the new rates as permanent rather than a launch promotion.
Price cuts are easy to report and easy to misread. A headline saying "50% cheaper" does not tell you what your own bill will do, because that depends on how much of your spend is input, how much is output, and how much of your input is cached. This post works through those three things with real numbers from our AI API Cost Calculator, so you can decide whether to switch this week or wait.
The new rate card
- GPT-6 Sol: $2 input, $10 output, $0.20 cached input, per million tokens.
- GPT-6 Luna: $0.10 input, $0.50 output, $0.01 cached input, per million tokens.
- GPT-6 Astra (unchanged, the flagship since early September): $10 input, $50 output, $1 cached input.
- Batch and Flex processing are half price. Fast mode is double.
- Prompts longer than 272,000 input tokens reprice the whole request: Sol moves to $4 / $15 and Luna to $0.20 / $0.75.
- Both models have a context window of about 1.05 million tokens and can produce up to 128,000 output tokens.
For comparison, GPT-5.6 Sol was listing at $4 / $20 and GPT-5.6 Luna at $0.20 / $1.20. One detail matters here: GPT-5.6 Sol's $4 / $20 was itself a promotional rate. Its standard price was $5 / $30. So if your budget was built on the standard GPT-5.6 Sol price, GPT-6 Sol is closer to 60% cheaper on input and 67% cheaper on output than the 50% in the headline.
What the cut does to a typical workload
Take a common setup: a support assistant or internal tool that sends about 5,000 input tokens per request (system prompt, a few retrieved documents, the user's question) and gets about 1,000 tokens back, 10,000 times a month. With no caching, on GPT-6 Sol:
Input
5,000 × $2.00 ÷ 1Mequals$0.01000
Output
1,000 × $10.00 ÷ 1Mequals$0.01000
Per month
$0.02000 × 10,000equals$200.00
That is $200 a month. The same workload was $400 on GPT-5.6 Sol at its promotional rate, $1,000 on GPT-6 Astra, and $10 on GPT-6 Luna. It also puts GPT-6 Sol at exactly the same per-token price as Claude Sonnet 5 ($2 / $10), and at half the per-token price of Claude Opus 5.5 ($4 / $20).
Caching changes the picture more than the price cut does
Most real prompts repeat. The system prompt, tool definitions and reference documents are often identical from one request to the next, and only the last few hundred tokens change. Cached input on GPT-6 Sol costs a tenth of the normal input rate. If 80% of each prompt is served from the cache, the same 10,000 requests cost:
Input (part cached)
(1,000 × $2.00 + 4,000 × $0.200) ÷ 1Mequals$0.00280
Output
1,000 × $10.00 ÷ 1Mequals$0.01000
Per month
$0.01280 × 10,000equals$128.00
$128 instead of $200. Notice where the money goes now: output tokens are $100 of the $128. Once caching is working, the output price is what decides your bill, and the fastest way to cut it further is to ask for shorter answers, not to switch models again. A prompt instruction such as "answer in at most five sentences unless asked for more" often saves more than any price change.
If the work does not need an instant answer (overnight classification, bulk summaries, back-filling tags on old records), Batch halves everything again. The same uncached workload through Batch comes to $100 a month on Sol.
Is Luna good enough to replace Sol?
At twenty times cheaper, Luna is tempting. For narrow, well-defined jobs it usually is enough: pulling fields out of an invoice, tagging support tickets, detecting language, short summaries of a single document. For a high-volume job with 2,000 input tokens and 500 output tokens, 50,000 times a month, the difference is $450 a month on Sol against $22.50 on Luna.
Where cheaper models tend to struggle is anything that has to be correct in the real world, not just look correct. Early independent testing after the launch found Luna producing configuration files that passed every automated format check but failed when actually deployed. That is the pattern to watch for with any low-cost model. If Luna's output is going straight into production (code, configuration, anything a customer sees without review), keep a real check in the loop: run the code, deploy the config to a test environment, or have Sol review Luna's work on a sample.
Three traps before you switch
- The 272,000-token cliff. The long-context surcharge applies to the entire request, not just the tokens above the line. A single request with 280,000 input tokens on Sol costs $4 per million on all of them, plus 1.5 times the output rate. If you feed whole codebases or long document bundles into one call, check how often you cross the line before estimating savings.
- Shorter answers are not always a saving. Independent reviewers noted that the new models tend to write somewhat shorter responses. That lowers your output bill, which is good for chat and coding, but if your product depends on long, complete documents (reports, contracts, detailed plans), check that nothing important is being skipped before you celebrate the cheaper invoice.
- Reasoning tokens are billed as output. Sol and Luna let you choose how much the model "thinks" before answering, including turning it off. Hidden reasoning is charged at the output rate, so a high reasoning setting on a simple task can quietly erase the price cut. For extraction and classification, start with reasoning off or low and only raise it if accuracy suffers.
Should you switch now?
If you are on GPT-5.6 Sol for everyday work, yes: change the model name, keep your prompts, and compare a week of output quality and cost. The saving is large and the risk is low. If you are on GPT-5.6 Sol at maximum reasoning for the hardest coding or computer-use tasks, test first, because early benchmark reads suggest the new model is cheaper per task rather than clearly better at the very top end. If you are on GPT-5.6 Luna, switching is almost free money, since output falls from $1.20 to $0.50 per million.
And if you are choosing between OpenAI and Anthropic, the per-token gap has mostly closed at the mid tier: GPT-6 Sol and Claude Sonnet 5 now list at the same $2 / $10. At that point the deciding factors are quality on your own tasks, how many tokens each model uses to finish them, and how well your prompts cache, not the rate card.
Tip: Put your own token counts into the AI API Cost Calculator and switch the model between GPT-6 Sol, GPT-6 Luna and your current model. Set the cache share to what your logs show, not a guess.
Questions people ask
- How much does GPT-6 Sol cost?
- $2 per million input tokens and $10 per million output tokens on the standard tier, with cached input at $0.20. Batch and Flex are half price.
- How much does GPT-6 Luna cost?
- $0.10 per million input tokens and $0.50 per million output tokens, with cached input at $0.01. It is OpenAI's lowest-cost GPT-6 model.
- Is the GPT-6 Sol price a limited-time promotion?
- OpenAI has described the GPT-6 Sol and Luna prices as permanent. That differs from GPT-5.6 Sol, whose $4 / $20 rate was a promotion over its $5 / $30 standard price.
- What happens above 272,000 input tokens?
- The whole request is repriced: GPT-6 Sol moves to $4 input and $15 output per million tokens, and GPT-6 Luna to $0.20 and $0.75.
- Is GPT-6 Sol cheaper than Claude?
- It lists at the same $2 / $10 as Claude Sonnet 5 and at half the per-token price of Claude Opus 5.5. The real cost difference depends on how many tokens each model uses for your task.
Sources
Get fee-change alerts
One short email when a price our calculators use changes. No other mail, unsubscribe anytime.