AI, Software

OpenAI launches Astra, its powerful (and controversial) new model

The new frontier model is particularly competent at agentic, computer-use tasks. OpenAI says GPT-6 Astra is 'the most intelligent and aligned model in the world'. News AI OpenAI says GPT-6 Astra is 'the most intelligent and aligned model in the world' The new frontier model is particularly competent at agentic, computer-use tasks.

GPT-6 Astra excels at computer-use and browsing, and handling tasks in the domains of software engineering, cybersecurity, science and general professional work, according to OpenAI. A flashy demo video the company released alongside the launch shows Astra handling everything from 3D modeling to building slideshows, and often taking care of multiple tasks across different domains at the same time (like ordering food while coding a game). Key to Astra's appeal is its ability to handle these multi-step workflows on your computer and in and out of your browser.

It's also allegedly able to do those complex tasks with "strong visual judgment," OpenAI says, and without drifting from its original directions or prompt. As with other model launches, the company has a new tranche of benchmarks to justify Astra's improved abilities, including a 98.6% score on ARC-AGI-3, one of the industry's benchmarks for measuring an AI's ability to solve unfamiliar problems. The score has some caveats, namely that AIs that have been put through the benchmark aren't all configured in the same way, and different system architecture (like whether they have persistent memory, for example) can produce different results.

Other improvements are more clear-cut: Astra scored 57.7% on Terminal Bench 4.0, a coding benchmark, and 59.3% on the Agent's Last Exam, a benchmark for measuring an AI's agentic capabilities, notably higher than GPT-5.6 Sol. OpenAI included even more benchmarks in its announcement, but in short, Astra is now at the top of most leaderboards in terms of performance. That raises the larger question of whether the AI model can be used and deployed responsibly.

The new model is stronger across the board, for better or worse. OpenAI launches GPT-6 Astra, and one of its founders thinks AGI is here.

OpenAI claims that Astra represents "a new frontier on computer and browser use," and that it handles tasks with unmatched "speed, accuracy, and safety.". OpenAI launches Astra, its powerful (and controversial) new model.

OpenAI launches GPT-6 Astra, and one of its founders thinks AGI is here The new model is stronger across the board, for better or worse. Astra is OpenAIโ€™s first model to reach its โ€œCriticalโ€ cybersecurity threshold and is more jailbreak-resistant than GPT-5.6 Sol.

OpenAI released Astra on Thursday, its latest AI model and โ€” according to the company โ€” its most powerful and capable one yet. OpenAI claims that Astra represents โ€œa new frontier on computer and browser use,โ€ and that it handles tasks with unmatched โ€œspeed, accuracy, and safety.โ€ The model is being made available Thursday to OpenAI customers that use Daybreak, its cybersecurity program.

OpenAI isnโ€™t officially calling Astra AGI, but its president says he personally thinks โ€œweโ€™re there.โ€ OpenAI launched GPT-6 Astra, its latest flagship AI model and what the company describes as the most capable model it has deployed to date. Astra brings notable gains in areas such as software engineering, computer use, and cybersecurity, and it will roll out to paid ChatGPT users and the API over the coming week.

Over the next week, it will also become available through OpenAIโ€™s paid plans โ€” including Pro, Plus, Enterprise, and Business accounts โ€” as well as through its API. In a call with journalists on Thursday, OpenAI president Greg Brockman said that Astra was the companyโ€™s โ€œmost intelligent and, also very importantly, our most aligned model yet.โ€ He added that it โ€œbrings together years of our research and big bets, with each breakthrough having built on the lastโ€ and that it represents a โ€œreal shift in what kind of work people can delegate to AI and how it can empower them.โ€ Much has been made about Astraโ€™s cyber capabilities.

And while OpenAI isnโ€™t officially calling this AGI, one of its founders now says he personally thinks weโ€™ve reached that point. OpenAI is particularly talking up Astraโ€™s ability to use computers and browsers, while president Greg Brockman described it as the companyโ€™s most intelligent model yet.

OpenAI published a blog earlier this week in which it discussed the modelโ€™s new capabilities, as well as new safeguards that have been instituted to make it a safer experience for users. The company said Thursday that it had tested Astra on a variety of security benchmarks to ensure its capabilities, and that โ€œIts ability to identify and develop zero-day exploits can help defenders find and patch weaknesses.โ€ The companyโ€™s focus on alignment โ€” that is, the tendency of a model to do what a user wants or is in their best interests โ€” canโ€™t help but seem like a response to the recent Hugging Face breach, in which an OpenAI agent escaped its sandboxed testing environment and hacked several companies (a very blatant example of misalignment).

The company also reckons Astra is its best model so far for software engineering, with testing showing improvements in areas like bug detection and working with codebases. A factor that many wonโ€™t find reassuring is that Astra is the first OpenAI model to hit the โ€œCriticalโ€ level for cybersecurity capabilities under the companyโ€™s Preparedness Framework.

OpenAI has also boasted about Astraโ€™s coding abilities, claiming that it is the โ€œbest model for software engineering to date.โ€ To back up that assertion, the company provides results from a variety of cyber-related benchmarking tests. Those tests seem to show that Astra scores higher than other existing models โ€” including OpenAIโ€™s own Sol and Anthropicโ€™s Fable โ€” when it comes to activities like finding bugs, executing terminal tasks, and answering queries about codebases.

OpenAI says that when given the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit well-protected systems without requiring a human to guide every step. Thatโ€™s potentially very useful for finding vulnerabilities before attackers do, but the downside probably doesnโ€™t need much explaining.

Astra is also possibly OpenAIโ€™s most controversial model yet due to its use of a particular reasoning technique known as opaque recurrence. This technique is known to obscure an important model-monitoring process known as chain of thought, which allows researchers to audit how and why an AI model made the decisions that it did.

OpenAI says Astra is also more resistant to jailbreak attempts than GPT-5.6 Sol, while its broader safety testing generally shows fewer signs of problematic behavior. Given Astraโ€™s stronger capabilities, the company is also applying additional monitoring to tool-based sessions as an extra layer of protection.

OpenAI has downplayed the degree to which Astra engages in opaque recurrence โ€” and on the call chief scientist Jakub Pachocki seemed to frame a certain amount of opacity as a natural outgrowth of model evolution. He stated that monitoring the reasoning process of a model was a critical form of oversight but that โ€œas model capabilities are increasing, monitorability is getting more challenging.โ€ He later added that one potential reason for this was that โ€œmore capable models can perform harder tasks using fewer language tokensโ€ or โ€œno language tokens,โ€ which he said then reduces the ability to monitor those particular tasks.


Discover more from ChuckysCarnage

Subscribe to get the latest posts sent to your email.

Leave a comment