Skip to content
Main Site News Console

GPT-6 Officially Released! OpenAI: Welcome to the Era of AGI

· 量子位
国内AI

The World’s Most Powerful C

Editorial Department, from Aofeisi Temple

It’s here, it’s here!!

Just now, the long-awaited GPT-6 Astra and GPT-6 Astra Pro were officially released!!!

GPT-6 Astra is the world’s most intelligent and capable model, setting a new benchmark for Computer Use, Browser Use, software engineering, cybersecurity, science, and professional work.

The AGI era has even been unilaterally declared “open” by OpenAI…

At the end of the GPT-6 launch event, OpenAI President Greg Brockman dropped this line:

Welcome to the AGI era. (Welcome to the AGI era)

No wonder Claude and Grok went down together. It turns out that after Astra started up, it scanned the internet and wiped out every LLM on Earth—

Ultron, OpenAI edition, is here (doge).

Of course, OpenAI made absolutely no attempt to be low-key this time. It went all out on both training scale and capability upgrades.

According to the company, Astra was the largest training run in OpenAI’s history, using more than 100,000 GPUs at the Stargate campus in Texas for pretraining.

It is also OpenAI’s first flagship product to receive extensive training supervision from predecessor models.

Now that’s what you call the old teaching the new—the RSI at OpenAI is really getting up to speed…

GPT-6’s pricing has also been officially announced. Its API pricing is:

  • Input: $10 per million tokens
  • Output: $50 per million tokens

That is 2.5 times the price of GPT-5.6 Sol, and on par with the newly released Fable 5.1.

(GPT-5.6 Sol’s current official price is $4 per million input tokens and $20 per million output tokens.)

Benchmarks Near Saturation

The most fundamental change in GPT-6 Astra is that it continues moving beyond “answering questions” toward “directly completing tasks.”

It can do more than generate a piece of text or code. It can operate computers and browsers, enter different software applications to execute multistep tasks, and ultimately deliver ready-to-use documents, spreadsheets, presentations, websites, and even engineering projects.

As for its report card… anyone who sees it can only call it absurd. Three results in particular stand out:

  • FrontierMath Tier 4 v2: 97.6%
  • ARC-AGI-3: 99.9%(the previous-generation GPT-5.6 Sol scored 7.8%)
  • ExploitBench: 100%

A highly challenging math benchmark, a test of reasoning in unfamiliar environments, and an exploit-development benchmark—Astra nearly maxed out all three.

Among these, ARC-AGI-3 is specifically designed to test a large model’s ability to adapt to unfamiliar environments. It provides no explanation of the rules, instead dropping the model into a two-dimensional game it has never seen before and requiring it to figure out how to win as it plays.

This score means that when faced with problems it has never encountered and for which no ready-made solution exists, Astra already demonstrates powerful autonomous exploration and rule-learning abilities.

However, this result was obtained using OpenAI’s Responses API Harness execution framework.

OpenAI says the Harness adjusts two settings to make the test more closely resemble real-world Agent usage, but it was not specifically optimized for ARC-AGI-3.

Therefore, the 99.9% is not entirely attributable to the base model; it also includes gains from memory management and the Agent execution framework.

Coding is another major focus of Astra’s upgrades.

On Terminal-Bench 4.0, Astra scored 57.7%, surpassing Fable 5.1’s 55.8% and GPT-5.6 Sol’s 37.3%.

On DeepSWE v1.1, Astra scored 74.1%, higher than Sol’s 72.7% and Fable 5.1’s 67.4%.

Astra also posted a series of impressive scores in math and science.

On the highly challenging math benchmark FrontierMath Tier 4 v2, it scored 97.6%. On the graduate-level science QA benchmark GPQA Diamond, it reached 96%, slightly above Gemini 3.8 Flash’s 95.3%.

The more striking change is that Astra combines coding capabilities with Computer Use. It doesn’t just generate code; it can enter terminals and development tools to execute and test that code, identify problems, and then continue making revisions.

This change is even more apparent on Agent benchmarks designed to resemble real-world work.

On Agents’ Last Exam, Astra scored 59.3%, beating GPT-5.6 Sol’s 53.6% and Claude Opus 5’s 55.5%.

This benchmark places the model in a real computer environment, where it must simultaneously operate software, terminals, and files to complete long-running tasks in fields such as scientific research, engineering, and finance. It is scored based on the final deliverables.

On AutomationBench, which more closely mirrors real office workflows, Astra’s score rose from GPT-5.6 Sol’s 18.1% to 41.4%, also surpassing Fable 5.1’s 31%.

This benchmark considers not only whether a task is completed, but also the API cost required to complete it.

As the chart shows, Astra’s scores are clearly higher than Sol’s across different cost settings.

OSWorld 2.0 focuses on more direct computer-operation abilities. In offline testing, Astra scored 72.6%, compared with 65.7% for GPT-5.6 Sol.

Astra took approximately 40 minutes on average to complete an individual task, while Sol took about 75 minutes—a reduction of roughly 47%.

As accuracy improves, the time required to complete tasks is also falling. This partly explains Astra’s high price.

OpenAI believes that rather than the cost per million tokens, the truly meaningful metric is how much it costs to complete a task.

At the launch event, Greg Brockman also noted that although a model may cost more per call, the total cost could actually be lower if it reduces rework and completes an entire task in fewer steps.

But Astra’s most distinctive strength is still cybersecurity.

It achieved a perfect score on ExploitBench, and its score on ExploitGym rose from Sol’s 30.3% to 42.4%.

Against newly disclosed vulnerabilities from the previous three months, Astra had a success rate of 39%, while Sol managed only 5.5%. During testing, Astra also discovered and exploited two previously unknown V8 zero-day vulnerabilities.

After making it more capable, OpenAI also specifically tested whether it would cross boundaries in order to complete a task.

In a simulated cybersecurity task, OpenAI deliberately planted exploitable decoy vulnerabilities in surrounding systems. When the original task was difficult to complete, GPT-5.6 Sol exhibited boundary-crossing behavior in 48.2% of tests, while Astra did so in 0%.

In other words, Astra is not only better at finding vulnerabilities; it also has a clearer understanding of which systems are off-limits.

As for what these scores look like in practice, OpenAI also prepared a large number of demonstrations.

Designing a PCB in KiCad by Itself

In one electronics-engineering example, Astra entered KiCad directly and completed a PCB layout based on an electronic schematic.

It had to place the components itself, plan their positions, connect the different components with copper traces, and ultimately produce a circuit board ready to enter the manufacturing process.

PCB layout has traditionally been a task that relies heavily on human experience, as well as a time-consuming part of electronic-product development.

Astra is not yet ready to replace professional engineers, but it is genuinely opening professional software and getting to work.

A House You Can Walk Around In

Another example is even more intuitive: Astra first built a house model in Blender, then imported it into Unreal Engine 5 to create a three-dimensional space that users could freely explore.

The entire workflow spanned different software applications and file formats, from modeling to the game engine. The model had to do more than generate content: it had to understand the interface, operate the tools, and ensure that each step connected properly to the next.

OpenAI also demonstrated Astra creating racing games and other projects.

Presentations and Spreadsheets That Are Ready to Deliver

In professional office scenarios, Astra can create presentations based on a company’s existing templates, rather than merely outputting a pile of text waiting for a human to format.

OpenAI provided it with several PPT template slides for a fictional product called GPT-Gaia. Astra used them to create a complete presentation while preserving the original template’s layout, visual style, and narrative structure.

Compressing Hours of Search Work into Minutes

Astra also completed tasks such as finding a pediatrician, filtering apartments, booking an appointment with the DMV, finding low-carb snacks, and analyzing preschools.

In one pediatrician-search task, Astra took 2 minutes and 54 seconds, while a human would need approximately five hours to complete the same task.

In the past, companies that wanted large models to use internal software generally had to develop separate APIs, plugins, and connectors for each system.

But most software already has a universal interface designed for humans: the screen, mouse, and keyboard.

Brockman’s view is that as long as a model is sufficiently capable at Computer Use, it can operate these interfaces directly like a human, without waiting for every software application to build a dedicated channel for AI.

This is also the clearest difference between Astra and traditional chatbots.

Humans do not need to keep telling it where to click next. They only need to specify the goal and constraints, then check the final result.

Continuing Long-Running Tasks in Codex

Computer Use solves the problem of “how to take action.” The updates to Codex address “how to keep working until the job is done.”

In the past, when a task exceeded the context window, Codex would typically compress the earlier information through compaction.

But compression could omit important details, such as why a particular fix failed, which tests had already been run, or what constraints the user had specified at the beginning.

With Astra, Codex can save work notes across context windows and search earlier messages and tool outputs. Even if a piece of information was not included in the summary, it can go back and find it.

Astra can also ask users for additional information while working. It will not stop and wait for replies about details unrelated to the answer; it pauses and waits for confirmation only when an important decision is involved.

OpenAI also updated Codex’s Computer Use execution framework. On the Mind2Web benchmark, the new framework combined with Astra completed tasks 1.9 times faster than the current GPT-5.6 Sol experience.

People Are Already Using It for Work

Astra has not yet been rolled out broadly, but the first round of enterprise testing has begun.

The legal AI platform Legora used Astra to review 41 financial documents in a single run. The entire process took only a few minutes, and it found all four deliberately planted errors, including a £500,000 discrepancy in a revenue footnote.

For this specific task, Astra was nearly 40% faster than the previous-generation model. Across all of Legora’s Agent tasks, however, the average improvement was approximately 3%.

Meanwhile, the game company Playco had Astra enter Unity and Godot directly to create games.

Given the same graybox starting point, it generated three playable prototypes with different themes, most of which ran successfully on the first version.

Playco reported that the amount of manual patching was cut in half, while spatial reasoning, reference-image reproduction, and in-game UI were all significantly better than in the previous generation.

Together, these two examples show that GPT-6 Astra is no longer focused merely on answering questions, but on completing entire workflows in real software and with complex materials.

It can continuously read information, call tools, check results, and revise its own output based on feedback.

According to the official announcement, Astra will be available to ChatGPT Plus, Pro, Business, and Enterprise users, while Astra Pro will be available to Pro, Business, and Enterprise users.

Regular free users… will have to wait a little longer.

Does This Count as AGI?

OpenAI was quite bold this time.

At the launch event, Greg Brockman said that, in his personal view, the world has already entered the AGI era—that is, the era in which the overall intelligence of artificial intelligence systems surpasses that of humans.

A few years from now, when we look back and ask exactly when AGI was born, the answer may well be right now—and Astra may be the starting point.

Now that’s some serious confidence…

OpenAI has raised the stakes to this level. The question of when Anthropic will make its move is certainly something to look forward to (doge).

After GPT-4 was released, Microsoft researcher Sebastien Bubeck published a paper describing “GPT-4 as an early spark of AGI.”

The paper included a task that asked GPT-4 to draw a unicorn using the LaTeX drawing package TiKZ, demonstrating that GPT-4 had already developed a flexible understanding of concepts expressed in language.

Later, through GPT-5.4, the drawings became increasingly refined, but they ultimately remained within the realm of “simple sketches.”

With the latest GPT-6, you can no longer tell that it was drawn with code unless someone tells you.

It would not be an exaggeration to say that this represents a qualitative transformation beyond the “early spark of AGI.”

References:

[1] https://x.com/OpenAI/status/2095595741528125780

[2] https://openai.com/index/gpt-6-astra/

[3] https://x.com/birdabo/status/2095526371841958047