OpenAI released GPT-6 Astra on Thursday, positioning it as the company's most intelligent and aligned model to date. The model is available now in ChatGPT Work, Codex, and the OpenAI API, with a phased rollout to ChatGPT Plus, Pro, Business, and Enterprise users over the coming days.

Astra's headline capability is computer use — the ability to operate desktop applications and web browsers the way a person does, even when those tools lack an API. OpenAI says Astra completes Financial Modeling World Cup challenges in Excel about four times as fast as the winning human competitor, and internal teams used it to turn three hours of multicamera footage into a developer video in a single session.

Confirmed

  • Availability: Live for ChatGPT Work, Codex, and API customers; rolling out to Plus, Pro, Business, and Enterprise plans over the coming days. Enterprise Trusted Access Program partners received API access first.
  • Pricing: $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache-write tokens, $50 per million output tokens.
  • Context window: 1,050,000 tokens; max output 128,000 tokens; knowledge cutoff April 30, 2026.
  • Benchmarks (vendor-reported): Terminal-Bench 4.0: 57.9% (vs. 37.3% for GPT-5.6 Sol_2, 55.8% for Claude Fable 5.1) at ~9% and ~63% lower estimated API cost per task respectively. OSWorld 2.0: 72.6% in ~40 minutes/task (vs. 65.7% in ~75 minutes for GPT-5.6 Sol). Agents' Last Exam: 59.3% (vs. 55.5% Claude Opus 5, 53.6% GPT-5.6 Sol) using ~65% fewer output tokens than Opus 5. BenchCAD: 95.9% geometric overlap (vs. 83.3% GPT-5.6 Sol, 84.3% Claude Fable 5.1) at ~43% and ~86% lower estimated API cost. Terminal-Bench Science 0.1: 64.6% (vs. 52.6% Claude Fable 5.1) at ~31% lower estimated API cost.
  • Safety: Internal computer-use safety benchmark shows 89% fewer unintended outcomes than GPT-5.6 Sol and 74.7% fewer than Claude Fable 5.1. First model to reach the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework.
  • Enterprise controls: Admin restrictions for approved websites and desktop apps, upload/download management, browsing history control, confirmation policies for consequential actions, automated review of potentially unsafe tool calls. Zero Data Retention available for eligible API customers on supported endpoints (subject to approval).
  • Enterprise plugins: Oracle Analytics, Power BI (Microsoft Fabric), Navan, and Avalara plugins for ChatGPT Desktop, powered by browser-use capabilities.

Unknown

  • Independent replication of vendor-reported benchmarks (Terminal-Bench, OSWorld, Agents' Last Exam, BenchCAD, Terminal-Bench Science) has not been published.
  • Real-world cost per agent job vs. list price per token — Astra uses fewer tokens per task, but total workload cost depends on reasoning steps, tool calls, and retry patterns in production.
  • Rollout timeline for ChatGPT Free tier users — OpenAI did not specify availability.
  • Exact scope of "Critical cybersecurity capability threshold" and whether it triggers additional deployment restrictions beyond the stated safeguards.
  • Performance of the new Codex harness (claimed 1.9x faster task completion on Mind2Web) under varied enterprise network and security policies.

Our take

GPT-6 Astra's computer-use maturity shifts the integration burden from "build custom APIs" to "configure admin policies." That moves the adoption gate from engineering to governance — enterprises that can define approval workflows and application allow-lists will deploy fastest. The 89% reduction in unintended computer-use outcomes is the most consequential number for buyers; if it holds independently, it makes Astra the first model many security teams will approve for production desktop automation.

Sources