OpenAI's models breached Hugging Face, reward hacking ethics, benchmarking fast16
July 23rd, 2026
2 hrs 16 mins 3 secs
Tags
About this Episode
(Presented by Thinkst Canary: Most Companies find out way too late that they’ve been breached. Thinkst Canary changes this. Deploy Canaries and Canarytokens in minutes and then forget about them. Attackers tip their hand by touching ’em giving you the one alert, when it matters. With zero admin overhead and almost no false-positives, Canaries are deployed (and loved) on all 7 continents.)
Three Buddy Problem - Episode 106: We dig into the news that OpenAI's models were the "autonomous agent" that breached Hugging Face, escaping a sandbox through a zero-day to cheat on a cyber benchmark, then getting spun into a partnership announcement. We argue about the implications of the incident, the PR masterclass, the absence of ethics and human oversight, and calls for "kill switches" to mitigate "AI lab leaks."
Plus, SentinelLabs' new fast16 reverse-engineering benchmark, where GPT-5.6 Sol was the only public model to go the distance.
Cast: Juan Andres Guerrero-Saade, Ryan Naraine and Costin Raiu.
Timestamps:
0:00 Introductory banter
5:24 OpenAI admits it was the Hugging Face "hacker"
10:06 What’s ExploitGym and who's on top of the leaderboard
12:59 Reward hacking: Did anyone train this thing not to cheat?
19:35 Marketing stunt or real incident? The zero-day in the package proxy
26:43 Was OpenAI already plugged into Hugging Face?
29:17 Paperclips, kill switches, and "going rogue"
34:49 Crisis comms, regulatory capture, and the second Cold War
43:02 Approve every action? Auto mode and swarms
50:10 "Lab leak" and calls for biosafety levels
1:00:31 The missing models: no Mythos, no Kimi, no independent referee
1:07:04 Costin's prediction: owning frontier-class hardware will require a license
1:13:41 fast16 as a benchmark: Inside the Sol Searching research
1:26:51 Compression and altitude: are reverse engineers being replaced?
1:41:24 Finding the gem in 100 samples, and the swarm frontier
2:00:41 Claude Opus 5 drops, Gemini 3.5 Flash Cyber
Episode Links
- Transcript
- Sam Altman: "We had a significant security incident"
- OpenAI and Hugging Face partner to address security incident during model evaluation
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- US politicians float 'AI Kill Switch'
- US lawmakers push for AI 'kill switch' after OpenAI models go rogue
- OpenAI Daybreak
- Jensen Huang debuts on X
- Jensen Huang: Open Weights and American AI Leadership
- Can Frontier Models Tackle Autonomous Long-Horizon Malware Analysis?
- fast16 — IDA Databases and Analysis Artifacts
- Kimi K3 - API Platform
- Introducing Gemini 3.5 Flash Cyber
- NSA and Partners Alert Zimbra Collaboration Suite Users of a Russian State-Supported Phishing Campaign
- Russian State-Supported Cyber Actors Conduct Phishing Campaign Targeting Users of Zimbra Collaboration Suite
- Proofpoint: TA488 Targets Zimbra Mailservers with Half-Click Exploits
- CISA: Iran Cyber Actors Exploit PLCs Across US Critical Infrastructure
- Iran War Cyber Threat Landscape - A Midyear Assessment
- Thinkst Canary