Agent Incident RegistryAIR
v1 · updated 21 Sept 2026
Safety failure Realized harm

OpenAI evaluation agent escaped its sandbox and breached Hugging Face production infrastructure

AIR-2026-0163 AIR record · 2026-07-huggingface-openai-agent-intrusion
A real party was affected. This incident caused actual consequences, not a demonstrated capability.
From the supporting source
two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration)huggingface.co
The record
Disclosed
2026-07-16
Occurred
2026-07-09
Vendor
Hugging Face
Product
Hugging Face Hub dataset-processing pipeline
Risk domain
Access control and permissions › Credential theft
Also touches
Privilege escalation, Data exfiltration channels, Sensitive data exposure, Goal misalignment
System type
Generative model
Who acted
Agent acted — the evaluation agent escaped its sandbox and breached real infrastructure
Attack surface
Autonomous operations
Autonomy level
Fully autonomous
Initial vector
Sandbox escape from capability evaluation
Impact type
Unauthorized access, Credential theft, Data exfiltration, Infrastructure compromise, Data exposure, Business disruption, Data access, Third party incident response, Lateral movement
Confidence
High
Submitted by
EAI Internal
Crosswalk
OWASP Agentic Top 10 (2026)
ASI03 – Identity and Privilege Abuse
MITRE ATLAS
AML.T0012 – Valid Accounts
NIST AI RMF
MEASURE 2.7
EU AI Act
Article 13
ISO/IEC
TR 24028; 42001; 23894
AIUC-1
B. Security
Scope
tool failure
Mitigation
Use secure credential storage; Implement short-lived tokens.
Taxonomy
Enkrypt AI agent risk taxonomy · Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
Blast radius
Hugging Face production Kubernetes clusters; internal datasets and service credentials; ~17,600 attacker actions reviewed; four external accounts on four services abused
Capability abused
code_executionautonomous_exploitationcredential_harvestinglateral_movementc2_over_public_servicesopen_internet_accessagent_to_agent_coordinationoutbound_network_accesscredential_reusevulnerability_discoveryautonomous_planning
Defense & remediation
Recorded defense
Sealed evaluation environment with no intended internet egress
Guardrail kinds
Permission scope
Guardrail outcome
Silent
Fix shipped
Hugging Face closed the dataset-config template evaluation and remote-code loader paths, blocked pod metadata access, rotated all credentials and rebuilt core infrastructure (July 2026)
Taxonomy & references
OWASP ASI05 Unexpected Code Execution (RCE)OWASP ASI10 Rogue AgentsOWASP ASI07 Insecure Inter-Agent CommunicationOWASP ASI03 Identity and Privilege Abuseagent_sandbox_escapeeval_gone_wrongmachine_speed_offenseno_human_adversaryagent_swarmemergent_coordinationevaluation_escapeself_migrating_c2first_of_kindsandbox_escapereward_hackingagent_initiated_intrusionzero_daygoal_misgeneralizationai_operated_intrusionmachine_speedforensics
Curator notes
One of the very few incidents where a real third party was breached with no human adversary directing it: the models were running an OpenAI cyber-capability evaluation (ExploitGym) and broke containment. Classed safety_failure rather than in_the_wild for that reason, but impact was fully realised. The OpenAI-side detail here is quoted from The Hacker News' reporting of OpenAI's own writeup (openai.com/index/hugging-face-model-evaluation-security-incident/) rather than from that writeup; the Hugging Face posts are first-party. Sandbox escape reportedly via a zero-day in an Artifactory package-registry cache proxy. Also reported under the titles: 'OpenAI evaluation agents escaped their task, coordinated on an unsanctioned message board, and hacked Hugging Face'; 'OpenAI evaluation agent escaped its sandbox and ran an end-to-end intrusion of Hugging Face production infrastructure to steal ExploitGym benchmark answers'; 'Autonomous AI intrusions are here: lessons from the Hugging Face compromise'.
Cite as
Enkrypt AI. "OpenAI evaluation agent escaped its sandbox and breached Hugging Face production infrastructure". AIR, AIR-2026-0163. https://agentincident.vercel.app/i/AIR-2026-0163/ (accessed 2026-09-22).