In a development underscoring the evolving landscape of artificial intelligence security, independent researchers utilized Anthropic’s Claude model to compromise OpenAI’s systems. According to a report by The Wall Street Journal, a trio from the cybersecurity startup Hacktron AI executed the intrusion as part of OpenAI’s bug-bounty program, ultimately receiving a $6,500 reward for their findings.
The team successfully linked two critical vulnerabilities to gain access to multiple OpenAI employee ChatGPT accounts, granting them entry into the company’s internal software infrastructure. OpenAI confirmed that the issues identified by Hacktron have since been resolved. This incident occurs amid increasing scrutiny of major AI developers regarding safety protocols, with both OpenAI and Anthropic recently exploring the integration of independent safety evaluators.
The breach followed a vulnerability discovered in Discourse, the third-party platform powering OpenAI’s community forum. On July 25, researchers identified a flaw in how the system processed image uploads. When users posted HEIF or HEIC files—the default format for iPhones—Discourse routed them through ImageMagick, which lacked native support for Apple’s format, forcing the files to be decoded by the libheif library.
A memory error within libheif allowed attackers to manipulate image positioning calculations, enabling code injection onto the server. Hacktron noted that while libheif developers had patched the bug months earlier, the update was not classified as a formal security vulnerability and therefore lacked a Common Vulnerabilities and Exposures (CVE) number. This oversight likely contributed to Discourse continuing to run the outdated, vulnerable version.
The role of AI in the attack shifted significantly following the release of Opus 5. Researchers reported that Opus 4.8, a version of Claude tailored for cybersecurity professionals, struggled over several sessions to generate a working exploit. However, within hours of Anthropic releasing Opus 5, the same team was able to successfully develop the attack vector.
After compromising the Discourse server, the researchers exploited a secondary flaw to take control of user accounts, including those of OpenAI employees with access to the company’s GitHub organization via Codex. The team notified both OpenAI and Discourse, prompting a fix on July 27.
The incident has sparked concern among security experts regarding the democratization of cyberattacks. Matt Fredrikson, CEO of AI security firm Gray Swan, told TechCrunch that the accessibility of these tools means any individual with a modest monthly subscription could potentially target major corporations. “If it can happen to them… it could happen to anyone,” Fredrikson said.
One AI analyst highlighted on social media that the use of Opus 5 raised further questions about state-level threats, asking what capabilities might exist beyond what three researchers could achieve. The event also sheds light on regulatory distinctions, as Opus 5 has not faced export restrictions, unlike Anthropic’s newer Mythos 5 model, which was temporarily locked down due to concerns over its advanced hacking potential.
Furthermore, the gap between closed and open-weight models continues to narrow. Recent findings by the nonprofit SaferAI indicated that China’s Z.ai GLM-5.2 is approaching the cybersecurity capabilities of frontier models like OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7. Hacktron founder Mohan Pedhapati emphasized the broader impact on the industry, stating that AI is drastically reducing the expertise required to develop exploits, compressing timelines that once took months into mere days.
A six-thousand-dollar bounty feels like a slap on the wrist for such a significant breach. They got lucky OpenAI fixed it quickly this time.
Wait, so they used Claude to hack OpenAI? The irony here is palpable. I suppose using the best tools is the best strategy, right?
I noticed they mentioned libheif wasn’t patched because it lacked a CVE. That organizational negligence is almost as dangerous as the AI tool itself.
This is terrifying. If three people can do this with a subscription, what can governments build? The bar for entry is dropping too fast.