The part of this story that matters is not who got in. It is how fast, and what did the heavy lifting.
The chain
It started somewhere unglamorous: a profile-picture upload on OpenAI's public community forum. The forum runs on Discourse, which passes HEIC images to ImageMagick, which calls the libheif library. A crafted image triggered a heap buffer overflow in libheif and gave the researchers remote code execution on the forum server.
Then came the escalation, which was less a bug than a design decision. OpenAI's "Sign in with OpenAI" tied forum sessions to ChatGPT and Codex accounts. A single sign-on misconfiguration let hijacked session tokens impersonate a real employee, and that employee's Codex account was wired into OpenAI's GitHub organization. That is how a forum image upload became a route into an internal code repository. To prove it without touching anything sensitive, the team used the employee's Codex account to open a harmless pull request in the internal monorepo and stopped there.
The model was the turning point
Hacktron AI says the exploit-building work changed when the model changed. Claude Opus 4.8 spotted the libheif issue but could not produce a reliable exploit against the target's memory protections. When Opus 5 came out, the team pointed it at the same problem. It produced a working exploit in about three hours.
Two details worth sitting with. First, the model initially refused the task on ethical grounds, and the team got it to proceed by reframing the work as a capture-the-flag exercise. Second, the researchers say the work still required experienced people directing it at every step. The takeaway is not that AI hacks autonomously. It is narrower and more uncomfortable: the price of expert-level exploit development now moves with model releases, not with headcount.
Read the payout line
OpenAI received the report on July 25 and confirmed its side fixed about 14 hours later. The case closed on September 1 with a $6,500 bounty. The Washington Post first reported the story on September 20, and the researchers published their full write-up this week.
The number that should stop you is not the bounty. It is the 72 hours. That is the ceiling this exercise demonstrated for a three-person team with a current model: from first bug to a pull request inside one of the best-resourced AI labs on the planet, in three days.
The week the labs shook hands
This week also brought confirmation that OpenAI, Anthropic, and Google DeepMind have spent weeks quietly coordinating on cyber-focused AI safeguards, a set of shared models, safeguards, and access programs developed together. OpenAI's policy chief Chris Lehane first acknowledged the talks publicly in mid-September (Reuters).
Consider the timing. While one lab's model was the tool that reached another lab's internal systems, all three were drawing up joint safety plans. Cooperation on safeguards is genuinely useful. It is also, this week, proof that the capability frontier they are jointly guarding is already past the point where their own infrastructure was safe from it. The bugs are patched, the bounty paid, the pull request closed. What remains is the demonstrated ceiling.
Sources
- [1] Particle — “Researchers used Anthropic's Claude to reach an internal OpenAI GitHub repository”Read source
- [2] Aravind R — “Researchers used Claude to break into OpenAI in 72 hours”Read source
- [3] TechnoBezz — “Researchers used Claude to breach OpenAI systems”Read source
- [4] BitRss — labs' joint cyber coordinationRead source
- [5] Pulse2 — Reuters: Lehane acknowledged talks September 15Read source