OpenAI published a customer story on September 14, 2026, describing how Perplexity uses GPT-6 Astra to write communications, edit real-world software, and monitor production systems with far less human oversight than earlier models. Perplexity cofounder and Chief Strategy Officer Johnny Ho said the model can craft communications, change software, and watch production systems in ways previous generations could not.

Ho highlighted testing as one of the most useful applications. With limited time for manual testing, he asks Astra to build a small testing program around an application. The model generates realistic responses that mimic external services such as a language model API or a connector, allowing Perplexity to check how its application responds and test workflows from start to finish. "We're actually able to trust it with full end-to-end systems and check in on it much less frequently than previous generations of models," Ho said.

Confirmed

  • Perplexity uses GPT-6 Astra for communications, software changes, and production monitoring.
  • Human check-in frequency is materially lower than with prior models, according to Perplexity.
  • Astra can stand in for external services during end-to-end testing, generating realistic API-like responses.
  • GPT-6 Astra launched in ChatGPT Work, Codex, and the API the week before the Perplexity story.
  • Pricing starts at $10 per million input tokens and $50 per million output tokens, per OpenAI.
  • On Terminal-Bench 4.0, Astra scores 57.9%, versus 37.3% for GPT-5.6 Sol_2 and 55.8% for Claude Fable 5.1, at roughly 9% and 63% lower estimated API cost per task respectively (OpenAI internal evaluation).
  • On OpenAI's internal computer-use safety benchmark, Astra produced unintended outcomes 89% less often than GPT-5.6 Sol and 74.7% less often than Claude Fable 5.1.
  • Enterprise admin controls can restrict approved websites, desktop applications, uploads, downloads, and browsing history; confirmation policies and automated review of tool calls are available.
  • Astra is the first model to reach the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework.
  • Zero Data Retention is available for eligible API customers on supported endpoints, subject to approval.

Unknown

  • Independent replication of Terminal-Bench 4.0 and safety benchmark results.
  • Real-world cost per agent job when workloads involve more reasoning steps or tool calls; list price per token is not total cost per task.
  • Rollout scope and SLA commitments for enterprise admin controls and Zero Data Retention.
  • Whether Perplexity's reduced check-in cadence translates to measurable reliability or incident-rate improvements in production.
  • Availability timeline for broader API access beyond ChatGPT Work and Codex.

Our take

Perplexity's endorsement signals that Astra's computer-use and reasoning gains are crossing a trust threshold for at least one AI-native company running live systems. The vendor-reported benchmarks are strong, but the missing independent replication and the gap between token pricing and total agent-job cost mean buyers should run their own workload profiling before committing budget at scale.

Sources