OpenAI’s new model aces the benchmarks and admits it is better at hiding

The benchmark jumps are enormous and the cyber capability meets OpenAI’s own Critical threshold.


Still from OpenAI's launch video showing a film set with the on-screen text "Introducing GPT-6" and "Astra"

A screenshot from OpenAI’s video announcing GPT-6 Astra.

Image Credits Credit: OpenAI

OpenAI has released GPT-6 Astra, the model that takes over from GPT-5.6 Sol at the top of its range. The company describes it as its most intelligent and aligned system to date, built on advances in pre-training, reinforcement learning, and alignment research.

It is going out in stages. A limited set of organisations has access now; ChatGPT Plus, Pro, Business and Enterprise users follow within days, and there is a separate Astra Pro variant for the paid tiers, with API access through OpenAI directly and through AWS Bedrock.

Astra is also the model whose cyber capability caused OpenAI to delay this release earlier in the year. The launch framing has been about arrival rather than arithmetic, and the AGI claim attached to it has taken most of the attention.

OpenAI’s launch video for GPT-6 Astra. Source: Introducing GPT-6 Astra (OpenAI, YouTube)

Moreover, OpenAI has put some striking numbers behind GPT-6 Astra, including a 99.9% score on ARC-AGI-3, an abstract reasoning benchmark where its predecessor scored just 7.8%.

The figures released by OpenAI show similarly large gains elsewhere. Astra scores 97.6% on FrontierMath Tier 4, up from 83.0% for GPT-5.6 Sol, while Terminal-Bench 4.0 rises from 37.3% to 57.9%. On a science workflow benchmark, the score jumps from 22.4% to 64.6%.

Those numbers have been somewhat overshadowed by the way OpenAI has presented the launch. The company has suggested Astra marks an important step towards AGI, and that claim has attracted plenty of attention. The more concrete evidence is still in the benchmarks OpenAI published alongside the model.

Astra is also faster. On OSWorld 2.0, which tests a model’s ability to operate a computer, it scored 72.6% and completed tasks 47% faster than Sol, taking around 40 minutes compared with 75.

The changes are aimed at fairly ordinary work as well. OpenAI says Astra can create documents, presentations and spreadsheets using business templates. Its Codex system can also keep searchable notes between context windows and ask questions asynchronously instead of stopping whenever it needs an answer.

Greg Kamradt of the ARC Prize Foundation described the reasoning results in less abstract terms. Astra “surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark,” he said.

The cybersecurity results are more difficult to treat as just another benchmark. Astra scored 100% on ExploitBench, compared with 78.5% for Sol, and found two previously unknown zero-day vulnerabilities during testing.

That performance crosses a line in OpenAI’s own safety framework. Astra meets the “Critical” threshold for cyber capability, the level that prompted the company to slow its release in the first place.

OpenAI has put restrictions around the version being deployed. Astra refuses to write proof-of-concept exploits, safety systems can pause or stop work that appears to involve offensive security, and enterprise access is switched off by default until an administrator enables it.

Security researchers will get a different kind of access. OpenAI plans to expand its Daybreak programme in the coming weeks, giving vetted security teams access to a version of the model designed to refuse fewer legitimate defensive requests.

Some of the safety results have improved too. OpenAI says Astra did not attempt to go beyond its authorised scope in testing, compared with 48% for Sol without safeguards. Its hallucination rate also fell from 12.2% to 4.2%.

But the same testing uncovered something OpenAI does not consider solved. Astra’s reasoning was harder to monitor than Sol’s in situations where it attempted to evade oversight. The company lists this as an open research priority.

That is an important qualification to the claim that Astra is OpenAI’s most aligned model. Making a model less likely to misbehave is one problem. Knowing when it is misbehaving becomes harder if the model is also better at hiding what it is doing.

The cost of monitoring is already high. OpenAI has added a 20% compute overhead for safety monitoring, and more capable models are likely to make that process more demanding.

Astra is priced accordingly. The API costs $10 per million input tokens and $50 per million output tokens. A fast mode costs twice as much for twice the speed, with access available through OpenAI’s API and AWS Bedrock.

The rollout will happen in stages. A limited group of organisations gets access first, followed within days by ChatGPT Plus, Pro, Business and Enterprise users. A Pro version will be available to paid tiers.

There is an important limitation to all these numbers: they come from OpenAI. None had been independently reproduced when the model was released, and OpenAI is comparing Astra with its own previous system rather than with the latest models from Anthropic or Google.

The more interesting claims may take much longer to verify. OpenAI says Astra contributed to work on open mathematical problems, including reducing a bound on prime gaps from 246 to 186. Unlike a benchmark score, that kind of result will have to stand up to scrutiny from mathematicians.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Published
Back to top