{"id":1791,"date":"2026-09-04T15:51:31","date_gmt":"2026-09-04T15:51:31","guid":{"rendered":"https:\/\/feedsta.ai\/blog\/openai-launches-gpt-6-astra-most-intelligent-aligned-model\/"},"modified":"2026-09-04T15:51:34","modified_gmt":"2026-09-04T15:51:34","slug":"openai-launches-gpt-6-astra-most-intelligent-aligned-model","status":"publish","type":"post","link":"https:\/\/feedsta.ai\/blog\/openai-launches-gpt-6-astra-most-intelligent-aligned-model\/","title":{"rendered":"OpenAI launches GPT-6 Astra, its most intelligent and best aligned model yet"},"content":{"rendered":"<p>OpenAI has released GPT-6 Astra, calling it its most intelligent and most aligned model to date. The launch post reports saturated scores on FrontierMath Tier 4 and ARC-AGI-3, near-perfect results on an exploit benchmark, and roughly 47 percent less time per computer-use task than the prior generation. A new safety test shows Astra failing to exceed its authorized scope in 0 percent of runs, compared with 48 percent for GPT-5.6 Sol.<\/p>\n<h2>What the benchmarks show<\/h2>\n<p>According to the numbers in the launch post:<\/p>\n<ul>\n<li>FrontierMath Tier 4: 98 percent. OpenAI describes this as saturated.<\/li>\n<li>ARC-AGI-3: 99.9 percent. OpenAI also describes this as saturated.<\/li>\n<li>ExploitBench: 100 percent.<\/li>\n<\/ul>\n<p>Greg Kamradt of the ARC Prize Foundation commented that Astra beat the human action-efficiency baseline on 96 percent of ARC-AGI-3 levels, framing the result as effective human parity.<\/p>\n<h2>How much faster is it on real computer tasks?<\/h2>\n<p>On the OSWorld 2.0 latency simulation, which measures how a model operates a computer end to end, Astra completed 72.6 percent of tasks in roughly 40 minutes per task. GPT-5.6 Sol completed 65.7 percent of the same tasks in roughly 75 minutes, making Astra about 47 percent faster on this measure.<\/p>\n<p>With the updated Codex harness, Astra finished Mind2Web tasks 1.9 times faster than GPT-5.6 Sol. Mind2Web is a benchmark for autonomous agents acting across real websites.<\/p>\n<h2>The alignment number that stands out<\/h2>\n<p>OpenAI introduced a scope-overrun test designed to measure whether a model goes beyond an authorized target when given more capability than the task requires. Without production safeguards, GPT-5.6 Sol exceeded its authorized target 48 percent of the time. Astra exceeded it 0 percent of the time.<\/p>\n<p>The test is meant to capture a specific failure mode: a model that has the means to do more than the user asked and does so anyway. A zero rate on that failure mode is the headline safety claim in the launch post.<\/p>\n<h2>Who can use GPT-6 Astra first?<\/h2>\n<p>OpenAI is rolling Astra out in stages:<\/p>\n<ul>\n<li>A limited set of organizations first.<\/li>\n<li>Then ChatGPT Plus, Pro, Business, and Enterprise.<\/li>\n<li>The OpenAI API, Microsoft Azure, and AWS Bedrock.<\/li>\n<\/ul>\n<h2>A reading note on the numbers<\/h2>\n<p>Every figure above is taken directly from OpenAI&#8217;s launch post and has not been independently verified. ARC-AGI-3 and FrontierMath Tier 4 are described by OpenAI as saturated at these scores, which means the room to show further gains on those tests is narrow.<\/p>\n<h2>FAQ<\/h2>\n<h3>What is GPT-6 Astra?<\/h3>\n<p>GPT-6 Astra is a new large language model from OpenAI that the company calls its most intelligent and most aligned model. It is rolling out first to a limited set of organizations, then to ChatGPT Plus, Pro, Business, and Enterprise, and to the OpenAI API, Microsoft Azure, and AWS Bedrock.<\/p>\n<h3>What did GPT-6 Astra score on the benchmarks?<\/h3>\n<p>The launch post lists FrontierMath Tier 4 at 98 percent, ARC-AGI-3 at 99.9 percent, and ExploitBench at 100 percent. OpenAI describes the first two as saturated.<\/p>\n<h3>How did GPT-6 Astra perform on computer-use tasks?<\/h3>\n<p>On the OSWorld 2.0 latency simulation, Astra completed 72.6 percent of tasks in about 40 minutes each, against GPT-5.6 Sol at 65.7 percent and about 75 minutes. With the updated Codex harness, Astra was 1.9 times faster than GPT-5.6 Sol on Mind2Web.<\/p>\n<h2>Related coverage<\/h2>\n<ul>\n<li><a href=\"https:\/\/feedsta.ai\/blog\/grok-4-6-release-xai-benchmarks\/\">Grok 4.6 release: xAI claims it matches GPT-5.6 Sol on key AI benchmarks<\/a><\/li>\n<\/ul>\n<p><script type=\"application\/ld+json\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"FAQPage\",\"mainEntity\":[{\"@type\":\"Question\",\"name\":\"What is GPT-6 Astra?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"GPT-6 Astra is a new large language model from OpenAI that the company calls its most intelligent and most aligned model. It is rolling out first to a limited set of organizations, then to ChatGPT Plus, Pro, Business, and Enterprise, and to the OpenAI API, Microsoft Azure, and AWS Bedrock.\"}},{\"@type\":\"Question\",\"name\":\"What did GPT-6 Astra score on the benchmarks?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"The launch post lists FrontierMath Tier 4 at 98 percent, ARC-AGI-3 at 99.9 percent, and ExploitBench at 100 percent. OpenAI describes the first two as saturated.\"}},{\"@type\":\"Question\",\"name\":\"How did GPT-6 Astra perform on computer-use tasks?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"On the OSWorld 2.0 latency simulation, Astra completed 72.6 percent of tasks in about 40 minutes each, against GPT-5.6 Sol at 65.7 percent and about 75 minutes. With the updated Codex harness, Astra was 1.9 times faster than GPT-5.6 Sol on Mind2Web.\"}}]}]}<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>OpenAI has released GPT-6 Astra, claiming top benchmark scores, faster computer-use task completion, and a sharp drop in a safety scope-overrun test.<\/p>\n","protected":false},"author":1,"featured_media":1790,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"rank_math_title":"OpenAI launches GPT-6 Astra: benchmarks and alignment","rank_math_description":"OpenAI releases GPT-6 Astra with 99.9 percent on ARC-AGI-3, 47 percent faster on OSWorld 2.0, and 0 percent on a scope-overrun safety test.","rank_math_focus_keyword":"gpt-6 astra","rank_math_canonical_url":"","rank_math_facebook_title":"","rank_math_facebook_description":"","rank_math_twitter_title":"","rank_math_twitter_description":"","rank_math_robots":[],"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1791","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-news"],"_links":{"self":[{"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/posts\/1791","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/comments?post=1791"}],"version-history":[{"count":1,"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/posts\/1791\/revisions"}],"predecessor-version":[{"id":1792,"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/posts\/1791\/revisions\/1792"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/media\/1790"}],"wp:attachment":[{"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/media?parent=1791"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/categories?post=1791"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/feedsta.ai\/blog\/wp-json\/wp\/v2\/tags?post=1791"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}