OpenAI expanded its GPT-6 model family on September 22, 2026, with two new releases — GPT-6 Sol and GPT-6 Luna — positioned as lower-cost alternatives to the flagship GPT-6 Astra model launched earlier this month. Both models cut API prices by 50% compared with their GPT-5.6 promotional rates, a reduction OpenAI attributes to improvements in caching and inference.
The company also published a series of vendor-run benchmark comparisons against Claude Opus 5 and Claude Fable 5.1, framing the release around cost-adjusted performance rather than raw capability alone.
Confirmed
- API pricing: GPT-6 Sol drops from $4 to $2 per million input tokens and from $20 to $10 per million output tokens versus GPT-5.6 Sol. GPT-6 Luna drops from $0.20 to $0.10 per million input tokens and from $1.20 to $0.50 per million output tokens versus GPT-5.6 Luna — a 50% cut on both models.
- Both models are available starting September 22 in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users. Free and Go users get GPT-6 Luna in the desktop app only. Neither model is yet available in the standard ChatGPT app. API identifiers are gpt-6-sol and gpt-6-luna.
- OpenAI improved prompt caching for GPT-6, offering a 90% discount on cached input-token reads, a new caching dashboard, a diagnostics tool for missed cache opportunities, and explicit cache breakpoints. OpenAI cites GitHub's report that these changes cut the share of prompt tokens requiring fresh processing by more than 50% across billions of requests.
- AutomationBench (business-workflow agent test, per Zapier's 1.0.6 benchmark): GPT-6 Sol at xhigh effort scored 33.2% at $0.27 per task — versus GPT-6 Astra at low effort (30.3% at 3.9x Sol's cost), Claude Opus 5 at max effort (26.9% at 11.1x Sol's cost), and Claude Fable 5.1 with Opus 5 fallback at max effort (31.4% at more than 8.9x Sol's cost, with fallback cost not reported).
- Agents' Last Exam: GPT-6 Sol at max effort scored 56.4%, above Claude Opus 5's highest reported score on the same evaluation, at 60% lower cost per task.
- Factuality (OpenAI's internal evaluation based on de-identified conversations where users flagged errors): GPT-6 Sol makes about half as many mistakes as its predecessor. GPT-6 Luna at higher effort matches GPT-5.6 Sol's factuality at roughly one-hundredth the cost.
- DeepSWE v1.1 (software-engineering tasks): GPT-6 Sol at max effort scored 68.8%, within 1.1 points of Claude Fable 5's highest reported score (69.9% at xhigh effort) at roughly 80% lower cost per task. GPT-6 Luna at max effort scored 66.6%, comparable to Claude Opus 5 and Fable 5 at medium effort, at 93% and 96% lower cost respectively.
- OSWorld 2.0 (computer-use tasks, offline set): GPT-6 Sol at xhigh effort scored 60.5% versus Claude Opus 5 at medium effort's 60.3%, at roughly 80% lower cost per task.
- OpenAI reports alignment improvements over the GPT-5.6 generation, including lower rates of misleading claims about coding work, with full results published in the model's system card.
Unknown
- All benchmark comparisons are vendor-reported — run by OpenAI in its own research environment or via its API. OpenAI's own footnote states that competitor scores were taken from publicly available reports rather than run independently, and the methodology or dates of those competitor runs are not specified.
- Real-world cost per agent task under sustained workloads: the reported per-task figures are list-price snapshots and do not account for cost scaling when reasoning effort or tool calls increase.
- Full ChatGPT app and web rollout timeline beyond OpenAI's statement that access would expand "gradually throughout the day."
- Whether GPT-6 Sol and Luna will receive their own system cards or continue to share GPT-6 Astra's.
Our take
OpenAI is using price as a lever to lock in developer workflows ahead of rivals matching its cost-intelligence curve. The caching discounts — which GitHub says already cut fresh token processing in half — make sustained agent workloads meaningfully cheaper for teams already building on OpenAI's stack. But every comparative number in this release is vendor-reported: OpenAI ran its own models and cites public reports, not independent runs, for competitors. Until third parties replicate the AutomationBench and DeepSWE results, the real gap against Anthropic's Opus 5.5 — released roughly 90 minutes before this announcement, per TechCrunch — stays an open question.