AI safety researchers say OpenAI’s Astra appears to do less of its reasoning in visible text, and OpenAI’s chief scientist has warned against a race into unmonitorability. He co-authored a 2025 position paper asking developers to evaluate and report exactly this, which the EU’s code of practice turns into a filing to the AI Office.
A debate has opened about whether OpenAI’s newest model can be watched while it works. Astra appears to do less of its reasoning in visible text, Semafor reported.
“It looks like it can solve hard competition math problems entirely in its head,” the AI safety researcher Ryan Greenblatt wrote on X. “This seems extremely concerning.”
Reports then surfaced that OpenAI may have deliberately limited visibility into the model’s outputs in order to improve its capabilities. The company has not confirmed that.
OpenAI’s chief scientist answered on Wednesday. “I want to prevent a race into unmonitorability kicked off by confused reporting,” Jakub Pachocki wrote, adding that he plans to write more on it.
TNW reported at launch that Astra’s reasoning became harder to monitor than its predecessor’s when the model tried to evade oversight, which OpenAI describes as an open research priority.
What makes the exchange unusual is that both sides of it wrote about this problem together, fourteen months ago, and agreed on what should happen next.
In July 2025 around forty researchers published a position paper calling chain-of-thought monitorability a new and fragile opportunity for AI safety.
The authors came from OpenAI, Google DeepMind, Anthropic, Meta, Amazon, the UK AI Security Institute and Redwood Research. Pachocki is one of them, and Redwood is where Greenblatt works.
The paper asked developers to build standardised evaluations for monitorability, to report the results, methodology and limitations in system cards, and to weigh monitorability alongside capability when deciding whether to train or deploy a model.
In Europe that reporting has a named recipient and a deadline, rather than a convention.
Signatories of the EU’s GPAI Code of Practice must submit a Model Report to the AI Office by market introduction, covering evaluations, mitigations and reports from external evaluators.
Each report must also contain five randomly selected samples of inputs and outputs from every relevant model evaluation, so that somebody outside the company can assess the work independently.
Inputs and outputs are precisely what remains when the reasoning stops arriving as text. OpenAI is a full signatory, and has already put a 20% compute cost on watching its own systems.
Get the TNW newsletter
Get the most important tech news in your inbox each week.