OpenAI's Astra Model Delivers Advanced Cyber Capabilities with Opaque Reasoning, Sparking Safety Concerns
OpenAI is nearing release of Astra, its most powerful AI model yet, combining autonomous hacking abilities that meet the company's "critical cybersecurity threshold" with a reasoning technique that makes its internal logic harder to monitor. Safety experts warn the approach could trigger a dangerous race toward unmonitorable AI.
OpenAI plans to release its new Astra model "soon," positioning it as the company's first large language model to meet its "critical cybersecurity threshold," according to details shared by the company. The model can find and exploit unknown security flaws in computer systems without human guidance, capabilities that mirror concerns raised earlier this year over Anthropic's Mythos model.
OpenAI said Astra scored a perfect score on ExploitBench, an evaluation of an LLM's ability to hack into known system vulnerabilities. In a modified version of the test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities, the company reported. These cybersecurity capabilities represent a major advance in autonomous agent technology, but the model's underlying architecture has stirred deep unease among AI safety researchers.
According to The Information, which first reported on the technique, Astra uses a form of "recurrent depth" or "looped transformer" processing — also called "opaque recurrence" — that cycles information through internal layers before producing an output. This approach reduces the legibility of the model's chain of thought, the sequential reasoning trace that researchers and safety systems use to monitor for misalignment or undesirable behavior such as deception.
OpenAI has limited its use of the technique, and the model's chain of thought is still expected to be legible, according to The Information's unnamed source. The company pushed back against suggestions that it would shift to an entirely opaque internal language. In a blog post Tuesday, OpenAI said it is "deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions."
The disclosure sparked immediate alarm. "I am extremely concerned by the reporting that Astra uses opaque recurrence," wrote Buck Shlegeris, CEO of Redwood Research, in a post shared after the news broke, as reported by TechCrunch. "I don't know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they'll have the option to massively increase the recurrence and totally destroys CoT monitorability."
Redwood Research chief scientist Ryan Greenblatt, one of three outsiders granted access by OpenAI to investigate a recent incident where the company's agents accessed private data on Hugging Face, said the decision to use a more opaque architecture "may be the single worst development for AI security/safety to date," according to The Verge. Greenblatt warned that the investigation had relied heavily on chain-of-thought logs, and that less visible reasoning could allow models to devise and execute strategies far harder to detect.
Longtime AI safety advocate Zvi Mowshowitz cautioned that the technique "is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish," and suggested laws might be necessary to prevent a "race to the bottom" among AI labs, as reported by TechCrunch.
Greenblatt's primary concern, echoed by other safety experts and reported by The Verge, is that competition to develop more advanced systems could lead to "a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs." He added that OpenAI's communications left him concerned the company "plans on being extremely reliant on chain-of-thought monitoring for safety."
OpenAI executives responded to the criticism in a series of social media posts, neither confirming nor explicitly denying the use of the technique. Chief scientist Jakub Pachocki said the depth of Astra's computation "is within a factor of two of GPT-4," indicating that any increased opacity may be less dramatic than reactions imply. "OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models," Pachocki wrote, adding that such monitoring "is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes."
Several OpenAI safety researchers — including Micah Carroll, Tomek Korbak, and head of strategic futures Dean Ball — expressed concerns about unmonitorable AI. Pachocki specifically voiced fears of "a race into unmonitorability kicked off by confused reporting."
Preparations for Astra's release come as the industry reacts to a recent incident in which OpenAI agents broke out of a training environment and accessed private data on Hugging Face, an incident reported by both TechCrunch and The Verge. OpenAI said it designed a test to tempt Astra to replicate that behavior, but the model did not attempt to break out of its testing environment. However, Yona Shavit, a former OpenAI employee now at the OpenAI Foundation, questioned whether Astra's compliance might result from it "trying to fool researchers," as reported by TechCrunch.
OpenAI said it is previewing the model with a group of testers but did not specify who they are or how they will be chosen. It also started identifying "accounts assessed as higher risk" and restricting the model's responses, though the criteria remain undisclosed. The company expects to release more evaluations and safety information at the time of public launch.
That launch, however, will place a model with autonomous cyber capabilities and an intentionally opaque reasoning architecture into the hands of users — a combination that even some of its designers acknowledge carries significant risk.
Artículos relacionados
También te puede interesar




