Skip to content
Main Site News Console

GPT‑6 Astra

· Simon Willison Translated
技巧LLM

GPT‑6 Astra

GPT-6 Astra “is rolling out to a small number of organizations starting today and will become available to all ChatGPT Plus, Pro, Business, and Enterprise users over the next few days. It will also be available through the OpenAI API and AWS.” I haven’t had a chance to try it myself yet, so I don’t have much more to say at this point.

Its API pricing matches that of Claude Fable 5 and 5.1: $10 per million input tokens and $50 per million output tokens. This is clearly OpenAI’s answer to Fable, and based on most of the benchmarks published by OpenAI itself, Astra appears to score higher.

Most impressive is Astra’s 99.9% score on the recently released ARC-AGI 3 benchmark, which was released in March this year. It is worth noting, however, that Fable 5 has not yet published results. In addition, the ARC-AGI blog post notes that the 99.9% score was achieved using OpenAI’s custom “Provider Adapter harness” at a cost of $19,000, while the model scored 62.7% using the default ARC-AGI harness at a cost of $26,000.

The Provider Adapter harness preserves opaque reasoning state between requests and uses a compaction mechanism to handle longer conversations, allowing the model to reuse work completed previously.

Given the recent Hugging Face incident, Astra’s strong performance on security tasks is not surprising. It scored 100% on ExploitBench (compared with 78.5% for GPT-5.6 Sol), 42.4% on ExploitGym (compared with 30.3% for Sol), and 99.2% within four attempts on the SRE-Bench binary reverse-engineering test, compared with 68.7% for Sol.

It is also stronger with long contexts: on OpenAI’s eight-needle benchmark, it scored 100% when processing 256K–512K tokens and 96.3% when processing 512K–1M tokens. OpenAI may have solved one of the longstanding challenges in long-context processing.

However, it does not outperform the competition in every area. Artificial Analysis notes that Astra still trails Fable on its Intelligence Index:

Comparable in intelligence to GPT-5.6 Sol: GPT-6 Astra scored 61 on the index, tying GPT-5.6 Sol. This is 5 points lower than Claude Fable 5.1’s highest score with fallback enabled. The model also trails Meta’s newly released Muse Spark 1.3 (highest score).

It performed better on their Coding Agent Index, however:

Leading the cost-efficiency frontier on the Coding Agent Index: At maximum effort, GPT-6 Astra costs roughly the same as GPT-5.6 Sol at maximum effort, but scores 2 points higher on the index. On a per-task basis, the model costs less than half as much as Claude Fable 5 for the same score.

I’ll write more about Astra once I gain access to it. After its release, the API model name will be gpt-6-astra.