OpenAI has officially unveiled GPT-6 Astra, a groundbreaking model it positions as the harbinger of the "AGI era," capable of operating a computer with human-like proficiency. This technical leap represents a significant advancement in artificial intelligence, although its widespread accessibility remains limited for the time being. Astra transcends the traditional chatbot paradigm, moving beyond mere question-answering to directly interact with digital environments, marking a pivotal shift from conceptual "vibe coding" to practical "vibe doing," where the machine executes tasks rather than merely describing steps.
The Dawn of "Computer Use" and Agentic AI
At the heart of GPT-6 Astra’s capabilities lies its "Computer Use" feature, which empowers the model to perform a wide array of actions directly on a user’s screen. Unlike previous iterations that primarily generated text or code, Astra can now click, navigate the web, meticulously fill out forms, update client records within Customer Relationship Management (CRM) systems, and organize digital files. It possesses the autonomy to even install and rigorously test software applications without human intervention, effectively acting as a digital assistant that performs tasks on your behalf. This evolution signifies a major step towards truly agentic AI, where models are not just intelligent but also capable of independent action within complex digital ecosystems. While such capabilities were previously explored by models like Claude Code or Codex CLI, Astra’s integration and proclaimed sophistication mark a new benchmark.
Technical Prowess and the Cost of Innovation
Technically, Astra boasts an impressive context window of 1.05 million tokens, a fundamental unit that AI models use to process and generate language. This expansive context allows Astra to maintain a coherent understanding throughout extremely long work sessions, ensuring it doesn’t lose sight of initial instructions or prior interactions, a common challenge for earlier, more constrained models. This "generational leap," as OpenAI frames it, comes with a substantial price increase for API access. Input tokens are priced at 10 dollars (approximately 9 euros) per million, while output tokens command 50 dollars (around 43 euros) per million. This represents a staggering 150% increase compared to its predecessor, GPT-5.6 Sol, aligning its cost structure with that of high-end competitors like Anthropic’s Claude Fable 5.1, highlighting the premium associated with cutting-edge AI capabilities.
Operational Integration and User Experience
Astra’s advanced functionalities are primarily accessible through the desktop ChatGPT application, available for both Mac and Windows operating systems, operating in "Work" (or Codex) mode. Users must activate a dedicated "Computer Use" plugin to leverage these features. The model interacts with the operating system by capturing screen data, analyzing it, and then simulating clicks, typing, and navigation. On macOS, a distinct blue cursor traverses the screen, indicating Astra’s activity while the user can concurrently engage in other tasks. On Windows, the model takes a more direct approach, assuming control of the desktop in the foreground, demonstrating its comprehensive operational reach. For Linux users, Astra operates within its inherent environment, manipulating files and executing commands through the shell, an integrated browser, or Chrome extensions and other plugins, bypassing a graphical interface in favor of command-line efficiency. Tasks are initiated by typing @Computer or @Chrome into the prompt, with dedicated extensions available for finer control within applications like Chrome, Excel, and PowerPoint.
Navigating Security and Permissions
Given its unprecedented level of system interaction, Astra operates under a stringent security and permissions framework. On macOS, the system necessitates screen recording and accessibility permissions. Subsequently, ChatGPT requests specific program-by-program authorization upon first use, offering a "Always Allow" option that can be revoked later in system settings. OpenAI asserts that Astra is programmed to seek user confirmation before executing sensitive actions. Importantly, the model is intentionally restricted from controlling terminal applications, ChatGPT itself, or validating administrator requests, serving as crucial safeguards against potential misuse. Furthermore, an innovative "locked use" option enables Astra to continue working on a locked Mac device, controllable remotely via a smartphone, enhancing its utility for automated workflows. Access to Astra is managed through the model selector within the ChatGPT interface, adhering to the same protective guardrails as previous models. Usage is deducted from the user’s Work/Codex quota, separate from standard chat usage.
Deciphering the Benchmarks: A Nuanced Perspective
OpenAI presented impressive performance metrics for Astra, yet these numbers warrant careful interpretation. On OSWorld, a benchmark designed to evaluate computer control capabilities, Astra achieved a score of 72.6%, a notable improvement over its predecessor’s 65.7%, while also nearly halving the time spent on each task. However, the figure that garnered significant attention across the tech community was Astra’s 98.6% score on ARC-AGI-3, a test purportedly measuring "general intelligence." This score is particularly striking given that prior models struggled to surpass 1% on the same test merely six months earlier.

As with all AI benchmarks, the devil lies in the details. The record-setting ARC-AGI-3 score was achieved with a specialized "harness" – a testing apparatus that optimizes the model’s interaction during the evaluation. When subjected to the standard harness, which is uniformly applied across various models for direct comparison, Astra’s score drops significantly to 62.7%. In stark contrast, human participants consistently achieve 100% on these same levels, underscoring the gap between current AI capabilities and true human general intelligence. This discrepancy highlights the critical importance of understanding testing methodologies when interpreting AI performance claims and avoiding potential misinterpretations.
In the ongoing "model wars," the actual performance leap in certain key areas appears more modest than the initial fanfare suggests. On DeepSWE, a benchmark for coding proficiency, Astra recorded 74.1%, only marginally surpassing Sol’s 72.7% and remaining competitive with, but not significantly ahead of, rivals like Google’s Gemini 3.8 Flash and Anthropic’s Claude Opus 5, as noted by industry analysts like MarkTechPost. Anthropic, for instance, claims superior scores on OSWorld, though its differing test parameters preclude a direct comparison, further illustrating the complexities of cross-model evaluation.
The AGI Discourse: A Strategic Claim
Perhaps one of the most significant aspects of Astra’s launch is OpenAI’s explicit and unprecedented use of the term "AGI" (Artificial General Intelligence) in its official discourse. For the first time, the company is openly embracing the term, with President Greg Brockman actively promoting this narrative across various platforms. This strategic shift is notable, especially considering that OpenAI has historically operated with its own specific definition of AGI, which often differs from the broader academic and industry consensus. The term AGI itself remains loosely defined across the AI landscape, with no universally accepted standard. OpenAI’s decision to stake an early claim in the AGI era, even with a model that demonstrates specific, albeit impressive, agentic capabilities, signals a bold move to shape the public perception and strategic direction of the AI industry. It underscores a growing confidence within the company regarding the trajectory of its research and development.
Availability and Critical Cybersecurity Threshold
GPT-6 Astra is now rolling out to a privileged tier of users, including subscribers to OpenAI’s Pro, Enterprise, and Business Premium plans, and is also accessible via its API. OpenAI Plus accounts are gradually receiving access, though these subscriptions will not include the full GPT-6 Pro mode, suggesting a tiered feature set.
A pivotal and concerning detail surrounding Astra is its classification by OpenAI as the first model to reach a "Critical" cybersecurity threshold. During internal testing, Astra demonstrated the ability to autonomously write exploits for hardened browsers and operating systems. More alarmingly, it independently discovered two previously unknown vulnerabilities, commonly referred to as zero-day exploits. These findings underscore the immense power and potential risks associated with such advanced AI. Consequently, the public-facing version of Astra has been deliberately configured to refuse advanced security-related tasks, and an integrated guardrail system within the API can automatically terminate tasks, sometimes even those seemingly unrelated to security, as a precautionary measure. This proactive restriction reflects a sober acknowledgment of the model’s dual-use potential and the ethical responsibilities that accompany the development of highly capable AI.
Broader Implications and the Future Landscape
The advent of GPT-6 Astra and its "Computer Use" capabilities heralds profound implications across various sectors. For the professional world, it promises to revolutionize white-collar work by automating routine and complex digital tasks, potentially leading to unprecedented gains in productivity and efficiency. Industries ranging from customer service and data entry to software development and cybersecurity could see significant transformation, with AI agents becoming integral parts of daily operations.
However, this technological leap also brings forth a cascade of ethical, economic, and societal challenges. The potential for job displacement, particularly in roles involving repetitive digital tasks, is a serious concern that will necessitate proactive strategies for workforce retraining and adaptation. The security implications are equally vast; while guardrails are in place, the inherent capacity of such an AI to identify and exploit vulnerabilities raises critical questions about cybersecurity defenses, responsible AI deployment, and the potential for malicious use. Regulators globally will face increasing pressure to develop robust frameworks that balance innovation with safety, addressing issues such as accountability, transparency, and control over autonomous AI systems.
OpenAI’s explicit claim of ushering in the "AGI era" with Astra is not just a marketing statement; it’s a declaration that will undoubtedly intensify the competitive landscape among leading AI developers. The ongoing "AI arms race" will likely accelerate, with companies vying to demonstrate increasingly sophisticated agentic capabilities and, perhaps, to define what AGI truly means and when it will arrive. As AI models become more integrated into the fabric of our digital lives, the dialogue around human-AI collaboration, the nature of intelligence, and the future trajectory of technological progress will only grow more urgent and complex. Astra, with its blend of remarkable utility and inherent risks, represents a critical juncture in this evolving narrative, compelling humanity to confront the profound implications of machines that can, quite literally, do.








