Aaron Grattafiori Explains How LLMs Hunt Patch Variants at Scale
October 9th, 2026
54 mins 36 secs
Tags
About this Episode
(Presented by TLPBLACK: A cybersecurity intelligence platform focused on sharing curated, high-sensitivity threat insights and research with trusted security professionals.)
Three Buddy Problem x Offensive AI Con: Umbriel AI's Aaron Grattafiori breaks down the new offensive-security playbook, using LLMs for patch variant analysis, why the vuln apocalypse hasn't become an exploit apocalypse, and how the death of security through obscurity puts closed-source binaries and firmware within reach.
Plus, we discuss rogue agents escaping sandboxes, the eval-driven loop behind recursive self-improvement, and the brutal cost asymmetry that makes offense cheap and defense almost unaffordable.
Cast: Juan Andres Guerrero-Saade, Ryan Naraine and Aaron Grattafiori.
Timestamps:
0:00 Introductory banter
1:12 Umbriel AI pitch: "Smart people in a place to cook”
3:08 Weird, exhausting, exciting time in security
4:00 What AI really changed: speed (and a 30-minute reimplementation)
6:36 Incomplete fixes and the variant-analysis pipeline
9:21 Uplifting, abstraction, and representing vulnerabilities
11:48 Experimentation and the new shape of security companies
14:46 The vuln apocalypse vs. the exploit apocalypse
16:36 Triage, verification, and reward hacking
18:08 Do models reintroduce bugs? Is vuln-free code possible?
24:22 Specialized models and the rise of Jev-style classifiers
26:47 Binaries, firmware, and the death of security through obscurity
31:15 Open source as a requirement and corporate contracting
33:54 Rogue agents: Hugging Face, OpenAI, and lab escapes
39:17 Why defense is so expensive, and running the rig
Episode Links
- Transcript
- Aaron Grattafiori on Twitter
- UmbrielAI on Twitter
- Aaron Grattafiori | LinkedIn
- Mark Dowd on the zero-day exploit marketplace
- Becca Lynch | NVIDIA AI Red Team
- Measuring LLMs’ ability to develop exploits (Anthropic)
- ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
- OpenAI: The Hugging Face incident and the road ahead
- What Is Jev? TypeSafe's System One Model Explained
- OpenAI Daybreak - Trusted Access for Cyber Overview
- Calif: Apple MIE Exploitation Challenge
- TLPBLACK