The rumors that have swirled through the technology industry for months have finally crystallized into a singular, consequential announcement: OpenAI has officially released GPT-6 Astra. This new frontier model represents a pivot in the company’s trajectory, with leadership declaring that the system likely marks the onset of artificial general intelligence (AGI). For OpenAI, this milestone fulfills a long-sought objective: the creation of highly autonomous systems capable of outperforming humans across a vast spectrum of economically valuable work.
During a private press briefing held earlier today, OpenAI co-founder and president Greg Brockman provided a remarkably direct assessment of the moment. Concluding the session with a statement that reverberated through the industry, Brockman declared, "Welcome to the AGI era."
While such a bold framing is rare even for the high-stakes world of frontier AI, the immediate, practical significance of Astra for enterprises is perhaps more tangible. OpenAI is positioning the model as the foundation for a new paradigm of computing—one where the traditional reliance on clicking through interfaces and typing on keyboards is no longer a prerequisite for digital labor. In its promotional materials, the company describes Astra as "the world’s best computer use model," designed to navigate software with the fluidity and intent of a human operator.
Unlike previous generations of AI that required developers to build custom API integrations for every specific application, Astra is designed to interact with the existing digital ecosystem. It can navigate web browsers, manipulate spreadsheets, manage desktop applications, and execute complex, multi-step workflows. Rather than simply providing instructions to a user, the model can generate finished documents, manage software installations, and perform administrative tasks independently.
To illustrate this shift, OpenAI released a promotional video that juxtaposed a nostalgic 1980s demo of an AI tasked with drawing a simple yellow circle with modern footage of its employees. In the contemporary demonstration, staff members interacted with Astra solely through voice commands. The model transitioned the requested yellow circle into a complex rocket ship, then into a fully realized 3D game, and finally executed an eBay listing—all without a single keystroke.
The rollout begins this Thursday for enterprise customers participating in OpenAI’s gated "Daybreak" program. In the coming days, the model will see a broader release, becoming available to ChatGPT Plus, Pro, Business, and Enterprise users, as well as via the OpenAI API and major cloud platforms, including AWS Bedrock and Microsoft Azure.
From Answering Questions to Operating Computers
The core of the enterprise value proposition for Astra lies in its capacity for computer use. OpenAI notes that the model can handle a wide variety of professional tasks: filling out online forms, updating customer relationship management (CRM) records, organizing calendars, conducting deep web research, and drafting professional communications. Beyond administrative support, it can manipulate complex spreadsheets, perform data analysis in Python notebooks, interact with business intelligence tools like Power BI, and even troubleshoot or install software.
This capability represents a significant departure from the standard generative AI architecture of the past few years. Until now, enterprises have been forced to act as middlemen, connecting models to corporate systems through a patchwork of APIs, plugins, and custom-built retrieval systems. During the briefing, Greg Brockman argued that computer-use agents could eventually bypass much of this friction. Because modern software is already designed with the human user—who uses a mouse and keyboard—in mind, an intelligent agent capable of mimicking those inputs can interface with any existing tool.
"We’ve been bottlenecked over this gigantic era by people writing connectors and very painstakingly building these connections into all these tools that people can already use," Brockman explained. With the advent of Astra, he suggested that an agent can simply "zip through spreadsheets, fill out forms, and navigate across web pages."
The development of this capability dates back to the earliest days of OpenAI, when researchers first discussed the potential of training an agent to interact with the basic building blocks of modern computing: pixels, keyboards, and mice. According to Brockman, the company has finally achieved a system that feels genuinely useful in that capacity. Benchmark testing supports this, with Astra scoring 72.6% on an offline subset of OSWorld 2.0 while completing tasks in 40 minutes, compared to its predecessor, GPT-5.6 Sol, which scored 65.7% and required 75 minutes per task—a 47% improvement in speed.
A Landmark in Model Training
Aidan Clark, an OpenAI researcher, described the development of Astra as the company’s largest-scale training run to date. It is the first model pretrained using more than 100,000 DBUs (Data-Based Units) at the company’s Stargate infrastructure. Furthermore, it represents a milestone in "recursive" training, as it is the first model for which previous iterations played a significant role in supervising the training process.
According to Clark, the leap in performance from Sol to Astra is quantitatively larger than the progression observed between any of the company’s previous models. The combination of massive-scale pretraining and reinforcement learning has allowed the model to connect disparate pieces of information and execute increasingly long, complex tasks.

The benchmark scores are striking across the board. Astra achieved 97.6% on FrontierMath Tier 4 v2, 74.1% on DeepSWE v1.1, 95.9% on BenchCAD, 96% on GPQA Diamond, and 100% on ExploitBench. Perhaps most notably, it reached 98.6% on the ARC-AGI-3 benchmark, a test specifically designed to measure an AI’s ability to generalize to unfamiliar problems rather than simply mimicking training data.
The Debate Over Defining AGI
The 98.6% score on ARC-AGI-3 has reignited a fierce debate within the AI community regarding what exactly constitutes "artificial general intelligence." While some see the high score as proof of AGI, others point to the role of the surrounding system—the memory, tools, and feedback loops—rather than the model itself.
In August, NVIDIA demonstrated that its Agentic Variation Operators (AVO) architecture could achieve a 100% score on the same benchmark. However, that result was achieved by pairing a base model, which on its own scored only 30%, with sophisticated mechanisms for memory and error recovery. This has led many to question whether "AGI" is a property of the foundation model or the entire deployed system.
For enterprises, however, the academic debate over "benchmark purity" may be secondary to operational reality. Companies are interested in outcomes: can the system reconcile accounts, investigate security incidents, or modify codebases reliably and cost-effectively? As Brockman acknowledged, the definition of AGI remains "fuzzy." Yet, he remains firm in his personal conviction: "I think it’s not unreasonable to feel that we are now in the AGI era."
A New Economic Lens for AI
With the introduction of Astra, OpenAI is also urging a shift in how enterprises calculate the cost of AI. The company argues that "price-per-token" is a poor metric for evaluating value, as different models and architectures use tokens in varying ways. Instead, OpenAI is pushing for "price-per-task" as the standard metric.
An inexpensive model that requires repeated human intervention, multiple retries, and constant oversight may end up being more costly than a premium model that executes a task correctly the first time. OpenAI’s data suggests that Astra’s most efficient configuration achieves better performance on DeepSWE v1.1 while reducing the estimated cost-per-task by approximately 57% compared to GPT-5.6 Sol.
Governance and the Challenge of Autonomy
The very autonomy that makes Astra useful also presents a significant governance challenge. Unlike a chatbot, which generates content for human review, an agent that operates a computer can directly manipulate data, files, and external systems. Consequently, OpenAI has prioritized safety and alignment in its development.
The company revealed that it paused certain frontier training runs for roughly two weeks to tighten security and infrastructure controls. This was not because Astra was deemed dangerous, but to ensure that the company’s safety, monitoring, and infrastructure protocols kept pace with the model’s rapid increase in capabilities.
OpenAI is implementing a "defense-in-depth" approach, combining built-in model refusals with system-level classifiers and offline detection to identify malicious patterns that might emerge over several steps. For high-risk users, the company can monitor broader conversational contexts to detect suspicious activity. In internal tests, when stripped of these production safeguards, older models frequently exceeded their authorized scope; Astra, by contrast, did not deviate from its authorized boundaries in any of the test cases.
Jakub Pachocki, OpenAI’s chief scientist, emphasized that progress in intelligence does not guarantee progress in alignment. He warned that as models become more capable, they become harder to monitor, making "observability" a critical infrastructure challenge. OpenAI is committed to pausing further scaling if it cannot maintain sufficient confidence in its ability to monitor and align these systems.
The Cybersecurity Frontier
Astra is also the first model to reach the "Critical" cybersecurity threshold under OpenAI’s Preparedness Framework. It has demonstrated the ability to discover previously unknown vulnerabilities and create exploit chains without human guidance. While these capabilities are dual-use, OpenAI is initially restricting the most advanced features to trusted defenders who protect critical digital infrastructure.
Ultimately, the arrival of Astra suggests that the transition to AGI may not be marked by a single, dramatic moment, but rather by a gradual economic shift. As organizations restructure their workflows to accommodate these autonomous agents—tasking them with complex objectives while retaining humans for high-level judgment and oversight—the presence of AGI may only become clear in retrospect. For now, the test for businesses is straightforward: how much consequential, independent work are they willing to entrust to the machine?

