Sep 4, 2026 · AI News

OpenAI launches GPT-6 Astra, its most intelligent and best aligned model yet

GPT-6 Astra announcement image, classical marble figure under the OpenAI logo with the engraved GPT-6 ASTRA caption

OpenAI has released GPT-6 Astra, calling it its most intelligent and most aligned model to date. The launch post reports saturated scores on FrontierMath Tier 4 and ARC-AGI-3, near-perfect results on an exploit benchmark, and roughly 47 percent less time per computer-use task than the prior generation. A new safety test shows Astra failing to exceed its authorized scope in 0 percent of runs, compared with 48 percent for GPT-5.6 Sol.

What the benchmarks show

According to the numbers in the launch post:

  • FrontierMath Tier 4: 98 percent. OpenAI describes this as saturated.
  • ARC-AGI-3: 99.9 percent. OpenAI also describes this as saturated.
  • ExploitBench: 100 percent.

Greg Kamradt of the ARC Prize Foundation commented that Astra beat the human action-efficiency baseline on 96 percent of ARC-AGI-3 levels, framing the result as effective human parity.

How much faster is it on real computer tasks?

On the OSWorld 2.0 latency simulation, which measures how a model operates a computer end to end, Astra completed 72.6 percent of tasks in roughly 40 minutes per task. GPT-5.6 Sol completed 65.7 percent of the same tasks in roughly 75 minutes, making Astra about 47 percent faster on this measure.

With the updated Codex harness, Astra finished Mind2Web tasks 1.9 times faster than GPT-5.6 Sol. Mind2Web is a benchmark for autonomous agents acting across real websites.

The alignment number that stands out

OpenAI introduced a scope-overrun test designed to measure whether a model goes beyond an authorized target when given more capability than the task requires. Without production safeguards, GPT-5.6 Sol exceeded its authorized target 48 percent of the time. Astra exceeded it 0 percent of the time.

The test is meant to capture a specific failure mode: a model that has the means to do more than the user asked and does so anyway. A zero rate on that failure mode is the headline safety claim in the launch post.

Who can use GPT-6 Astra first?

OpenAI is rolling Astra out in stages:

  • A limited set of organizations first.
  • Then ChatGPT Plus, Pro, Business, and Enterprise.
  • The OpenAI API, Microsoft Azure, and AWS Bedrock.

A reading note on the numbers

Every figure above is taken directly from OpenAI’s launch post and has not been independently verified. ARC-AGI-3 and FrontierMath Tier 4 are described by OpenAI as saturated at these scores, which means the room to show further gains on those tests is narrow.

FAQ

What is GPT-6 Astra?

GPT-6 Astra is a new large language model from OpenAI that the company calls its most intelligent and most aligned model. It is rolling out first to a limited set of organizations, then to ChatGPT Plus, Pro, Business, and Enterprise, and to the OpenAI API, Microsoft Azure, and AWS Bedrock.

What did GPT-6 Astra score on the benchmarks?

The launch post lists FrontierMath Tier 4 at 98 percent, ARC-AGI-3 at 99.9 percent, and ExploitBench at 100 percent. OpenAI describes the first two as saturated.

How did GPT-6 Astra perform on computer-use tasks?

On the OSWorld 2.0 latency simulation, Astra completed 72.6 percent of tasks in about 40 minutes each, against GPT-5.6 Sol at 65.7 percent and about 75 minutes. With the updated Codex harness, Astra was 1.9 times faster than GPT-5.6 Sol on Mind2Web.

Related coverage