Researchers Used Anthropic's Claude to Breach OpenAI's Systems, Uncovering Security Flaws
A three-person security team at Hacktron AI exploited a vulnerability in OpenAI's community forum using Anthropic's Claude Opus 5, gaining access to employee accounts and internal code repositories before reporting the flaws for a $6,500 bug bounty.
Independent security researchers used Anthropic's Claude to break into OpenAI, exposing gaps in the ChatGPT-maker's defenses, according to reports from The Wall Street Journal and published by TechCrunch and The Verge on Thursday.
A three-person team at startup Hacktron AI carried out the attack as part of an OpenAI bug-bounty program. According to TechCrunch, OpenAI paid the startup a $6,500 award after Hacktron reported its findings. OpenAI said it has resolved the issues uncovered.
The researchers found a path into OpenAI on July 25 via a flaw in Discourse, the third-party software powering OpenAI's community forum, TechCrunch reported. The entry point was a mundane image upload. When users posted HEIF or HEIC image files (the default iPhone format) to the forum, Discourse passed them through a chain of tools to convert them into standard JPEGs.
The file first went to ImageMagick, a decades-old open-source utility for resizing images. Because ImageMagick cannot handle Apple's format, it handed the file to a library called libheif for decoding. A memory bug inside libheif exposed a path for an attacker to inject their own instructions. Feeding the library a specially crafted image caused it to miscalculate image positioning, which hijacked the server.
TechCrunch reported that the bug had been fixed months earlier by libheif's developers, but the fix was never formally flagged as a vulnerability. It never received a CVE (common vulnerabilities and exposures) number, the industry standard for tracking known security weaknesses. Hacktron said that may explain why the software used by Discourse was still running the vulnerable version.
The Claude model used — a special version of Opus 4.8 made available for cybersecurity researchers — could not build a working exploit at first, TechCrunch reported. That changed overnight when Anthropic released Opus 5. "Opus 4.8 struggled across several sessions to produce a working exploit," Hacktron wrote in a blog post, as reported by TechCrunch. "Within hours of Opus 5’s release, we gave it the same problem and it succeeded."
Once inside the Discourse server, the researchers found another flaw that let them take over users' ChatGPT and Codex accounts, including those belonging to OpenAI employees, TechCrunch reported. According to The Verge, the researchers gained access to OpenAI's GitHub repository, called "Monorepo," which contains "OpenAI's algorithmic secrets," citing The Wall Street Journal. They stopped short of accessing internal code themselves but sent a pull request from an employee's Codex account to prove they had access.
Hacktron says it alerted OpenAI and Discourse, which issued a fix on July 27, per TechCrunch. "We then took over an OpenAI employee’s account, whose Codex was connected to OpenAI’s GitHub organization," Hacktron wrote in a summary of the event, as reported by TechCrunch.
The Verge reported that Hacktron's "HEIF Heist" project took only one or two days to adapt to different companies including OpenAI, Slack, Meta, GitHub Ent, Rails, Next.js, ImageMagick, and others, using less than $3,000 in tokens. To the researchers' knowledge, only one target, Shopify, detected the activity.
The incident highlights how off-the-shelf AI tools can find vulnerabilities in even advanced companies' infrastructure. "For $200 a month, anyone can use these tools and hack into a company like OpenAI," Matt Fredrikson, CEO of AI security firm Gray Swan, told TechCrunch. "If it can happen to them — and I don’t think they’ve been slouching recently on cybersecurity hygiene — it could happen to anyone."
Hacktron founder Mohan Pedhapati noted on social media, as reported by TechCrunch: "AI is reducing the amount of scarce expertise needed to develop exploits. Work that once took months can now take days."
According to Ars Technica, OpenAI said in a statement: "We thank the researchers for contacting us and sharing their findings," adding that it had fixed the issues. Anthropic declined to comment.
The disclosure came as Anthropic published data showing rapid increase in AI use to develop new models, Ars Technica reported. Anthropic said 26 percent of research and development work was "led by" its Claude model, up from 1 percent in March, meaning AI completed the majority of tasks based on human instruction and under supervision. The company said it shared the data to help the public "understand how close the world is to reaching recursive self-improvement," the point at which AI can train and improve itself or new models. Its models did not yet operate fully autonomously for any research studied, Anthropic added.
Hacktron did not immediately respond to a request for comment, per Ars Technica.
Related articles
You might also like




