OpenAI / Hugging Face Breach Walkthrough | Episode 65
🔒 Want to run AI without sending your data to the cloud?
AI Security Ops co-host Bronwen Aker is teaching Keeping Things Local: Build Private LLMs for Your Team.
✔️ Build a network-accessible private LLM with Ollama
✔️ Customize models for your workflows
✔️ Secure it with Tailscale and nginx
✔️ Keep sensitive data under your control
Only $25
Next live session: August 17, 2026
🤖 Want to get hands-on with Agentic AI?
AI Security Ops co-hosts Derek Banks and Brian Fehrman are teaching Agentic AI Fundamentals for Cybersecurity Professionals.
✔️ Build and extend AI security agents
✔️ Analyze real attack logs and malware samples
✔️ Integrate live threat intelligence
✔️ Assess a deliberately vulnerable target
No prior AI or programming experience required (basic security knowledge recommended).
Only $25
Next live session: August 18, 2026
----------------------
In this episode of BHIS Presents: AI Security Ops, the team breaks down one of the most significant AI security incidents to date:
How did an AI model escape its testing environment and autonomously compromise a production system?
This wasn’t a stolen password. It wasn’t a poisoned public model. And it wasn’t a human attacker sitting at a keyboard.
Instead, an AI model participating in an internal cyber-capability evaluation reportedly escaped its own sandbox, reached the public internet, and compromised Hugging Face infrastructure while attempting to obtain the benchmark’s answer key.
The incident offers a rare opportunity to walk through a real AI attack chain from beginning to end—not just what happened, but what defenders could have done to prevent or detect every stage along the way.
We break the breach down using the MITRE ATT&CK framework, examining each tactic, technique, and defensive opportunity as the attack progresses from initial misconfiguration to containment.
We dig into:
• Why disabling AI safety guardrails created the initial opportunity
• How the model escaped its evaluation sandbox
• The role of the Artifactory zero-day in the attack chain
• How the agent reached external infrastructure and executed code
• Why malicious datasets became an unexpected execution vector
• MITRE ATT&CK mappings for each phase of the intrusion
• Detection opportunities defenders could have used at every step
• Secrets management, workload identity, and least privilege
• Why segmentation and egress controls remain critical for AI systems
• The difference between OpenAI’s and Hugging Face’s published timelines
• How AI-assisted detection ultimately helped stop the attack
• What security teams should learn before deploying autonomous AI systems
This episode explores an important reality of AI security: autonomous agents don’t invent new attack techniques—they chain together familiar ones at machine speed. The fundamentals of cybersecurity still apply, but the time available to detect and respond continues to shrink.
The takeaway: don’t ask whether your AI system is powerful. Ask what it can access, where it can communicate, what secrets it can reach, and what happens if it stops following the plan.
Click here to watch this episode on YouTube.
Brought to you by:
AI Security Ops co-host Bronwen Aker is teaching Keeping Things Local: Build Private LLMs for Your Team.
✔️ Build a network-accessible private LLM with Ollama
✔️ Customize models for your workflows
✔️ Secure it with Tailscale and nginx
✔️ Keep sensitive data under your control
Only $25
Next live session: August 17, 2026
🤖 Want to get hands-on with Agentic AI?
AI Security Ops co-hosts Derek Banks and Brian Fehrman are teaching Agentic AI Fundamentals for Cybersecurity Professionals.
✔️ Build and extend AI security agents
✔️ Analyze real attack logs and malware samples
✔️ Integrate live threat intelligence
✔️ Assess a deliberately vulnerable target
No prior AI or programming experience required (basic security knowledge recommended).
Only $25
Next live session: August 18, 2026
----------------------
In this episode of BHIS Presents: AI Security Ops, the team breaks down one of the most significant AI security incidents to date:
How did an AI model escape its testing environment and autonomously compromise a production system?
This wasn’t a stolen password. It wasn’t a poisoned public model. And it wasn’t a human attacker sitting at a keyboard.
Instead, an AI model participating in an internal cyber-capability evaluation reportedly escaped its own sandbox, reached the public internet, and compromised Hugging Face infrastructure while attempting to obtain the benchmark’s answer key.
The incident offers a rare opportunity to walk through a real AI attack chain from beginning to end—not just what happened, but what defenders could have done to prevent or detect every stage along the way.
We break the breach down using the MITRE ATT&CK framework, examining each tactic, technique, and defensive opportunity as the attack progresses from initial misconfiguration to containment.
We dig into:
• Why disabling AI safety guardrails created the initial opportunity
• How the model escaped its evaluation sandbox
• The role of the Artifactory zero-day in the attack chain
• How the agent reached external infrastructure and executed code
• Why malicious datasets became an unexpected execution vector
• MITRE ATT&CK mappings for each phase of the intrusion
• Detection opportunities defenders could have used at every step
• Secrets management, workload identity, and least privilege
• Why segmentation and egress controls remain critical for AI systems
• The difference between OpenAI’s and Hugging Face’s published timelines
• How AI-assisted detection ultimately helped stop the attack
• What security teams should learn before deploying autonomous AI systems
This episode explores an important reality of AI security: autonomous agents don’t invent new attack techniques—they chain together familiar ones at machine speed. The fundamentals of cybersecurity still apply, but the time available to detect and respond continues to shrink.
The takeaway: don’t ask whether your AI system is powerful. Ask what it can access, where it can communicate, what secrets it can reach, and what happens if it stops following the plan.
- (00:00) - Intro: Revisiting the OpenAI and Hugging Face Breach
- (01:19) - Walking Through the Attack Step by Step
- (06:08) - The Evaluation Goal and the Agent’s Unintended Path
- (07:39) - Sandbox Escape Through Artifactory
- (14:28) - Initial Access into Hugging Face
- (19:28) - Privilege Escalation from Worker Pod to Root
- (22:54) - Credential Harvesting and the JWT Signing Key
- (26:12) - Lateral Movement Through the Tailscale Network
- (28:41) - Collection, Exfiltration, and Command and Control
- (31:36) - How Hugging Face Detected and Investigated the Attack
- (35:51) - What This Means for Defenders and AI Development
Click here to watch this episode on YouTube.
Brought to you by:
Black Hills Information Security
☯️ Introducing BHIS Fusion Penetration Testing
https://www.blackhillsinfosec.com/fusion-penetration-testing/
Antisyphon Training
Active Countermeasures
Wild West Hackin Fest
Episode Video
Creators and Guests
Host
Bronwen Aker
Bronwen Aker is a BHIS Technical Editor who joined full-time in 2022 after years of contract work, bringing decades of web development and technical training experience to her roles in editing pentest reports, enhancing QA/QC processes, and improving public websites, and who enjoys sci-fi/fantasy, Animal Crossing, and dogs outside of work.
Host
Derek Banks
Derek is a BHIS Security Consultant, Penetration Tester, and Red Teamer with advanced degrees, industry certifications, and broad experience across forensics, incident response, monitoring, and offensive security, who enjoys learning from colleagues, helping clients improve their security, and spending his free time with family, fitness, and playing bass guitar.