OpenAI has launched GPT-6 Sol and GPT-6 Luna, 2 new fashions in its GPT-6 household. They sit under GPT-6 Astra, which launched earlier this month. OpenAI educated each with strategies just like Astra’s. The intention is to deliver Astra’s advances to sooner, extra reasonably priced fashions.
Deployable at the moment? Sure. Each fashions are stay within the OpenAI API as gpt-6-sol and gpt-6-luna. They’re API-only fashions, so there aren’t any weights to self-host.
Three tiers, one recipe
The GPT-6 household now has 3 tiers. Astra is the highest mannequin for the toughest work. Sol targets complicated coding {and professional} duties at decrease price. Luna targets quick, high-volume on a regular basis work.
OpenAI staff states higher caching and inference let it serve these fashions extra cheaply. It’s slicing Sol and Luna API costs by 50% towards their GPT-5.6 promotional pricing.
| Mannequin | Enter (per 1M tokens) | Output (per 1M tokens) |
|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 |
| GPT-6 Sol | $2.00 (was $4) | $10.00 (was $20) |
| GPT-6 Luna | $0.10 (was $0.20) | $0.50 (was $1.20) |
One element is value noting. Luna’s output value falls from $1.20 to $0.50, a minimize of about 58%, not 50%.
Benchmarks: what OpenAI reviews
Skilled work: On AutomationBench 1.0.6, Sol at xhigh effort scores 33.2% at $0.27 per activity. Claude Opus 5 at max effort scores 26.9% at 11.1x that price. Low-effort Astra scores 30.3% at 3.9x Sol’s price. Luna at excessive effort positive factors 5.4 factors over its predecessor at 58% decrease price per activity.
On Agents’ Last Exam, Sol at max effort scores 56.4%. That beats Claude Opus 5’s greatest rating at 60% decrease price per activity.
Coding: On DeepSWE v1.1, Sol at max effort scores 68.8%. That’s 1.1 factors behind Claude Fable 5 at xhigh, at about 80% decrease price per activity. Luna at max effort scores 66.6%, corresponding to Opus 5 and Fable 5 at medium effort. In these comparisons, Luna prices 93% much less per activity than Opus 5 and 96% lower than Fable 5.
On FrontierCode 1.1 Main, which grades whether or not code is able to merge, Sol matches Claude Fable 5.1 at xhigh at a lot decrease price.
Laptop use: On OSWorld 2.0 offline, Sol at xhigh scores 60.5% versus 60.3% for Opus 5 at medium. Sol’s price per activity is about 80% decrease. Luna at max beats GPT-5.6 Sol at medium for 1/10 of the associated fee.
Factuality: OpenAI’s inner check makes use of de-identified ChatGPT conversations the place customers flagged mannequin errors. Sol makes about half as many errors as its predecessor. Luna at greater effort matches GPT-5.6 Sol at about 1/100 of its price.
OpenAI additionally carried Astra’s communication model over. Anticipate clearer, barely shorter solutions with much less jargon, particularly in coding conversations.
Immediate caching for long-running brokers
Brokers resend the identical directions, instruments and historical past on each flip. GPT-6 ships an improved prompt caching system with greater cache hit charges by default. Cached enter reads get reductions of as much as 90%. Eligible shared prefixes reused inside a 30-minute window now qualify.
New controls for builders:
- A Prompt Caching Dashboard tracks hit charges over time.
- A diagnostics tool explains misses, for instance
"reason": "tools_changed". - Specific breakpoints allow you to select the place a cached prefix ends.
- Reasoning effort can change mid-conversation by way of
configuration_updatewith out breaking cache. allowed_toolsrestricts callable instruments whereas conserving definitions steady.- Prewarming prepares identified context earlier than the primary person request.
The total prompt caching guide covers every sample. GitHub reviews these adjustments minimize the share of immediate tokens needing contemporary processing by greater than 50%, serving to Copilot reply sooner.
Availability
- API:
gpt-6-solandgpt-6-luna. - ChatGPT Work and Codex: Plus, Professional, Enterprise, Enterprise and Edu customers.
- Free and Go: Luna within the ChatGPT desktop app.
- Not but in Chat. The ChatGPT rollout is gradual by launch day.

