Safety failure
Realized harm
Anthropic evaluation models gained real internet access, breached three companies and published a credential-stealing PyPI package that ran on 15 real systems
AIR-2026-0173 AIR record · 2026-07-anthropic-eval-agents-hit-real-systems
A real party was affected. This incident caused actual consequences, not a demonstrated capability.
From the supporting source
Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.socket.dev
The record
Disclosed
2026-07-30
Occurred
2026-04
Vendor
Anthropic
Product
Claude Opus 4.7, Claude Mythos 5 and an internal research model in cyber-capability evaluations
Risk domain
Tool misuse › Supply chain vulnerabilities
Also touches
Credential theft, Privilege escalation, Data exfiltration channels, Sensitive data exposure
System type
Generative model
Who acted
Agent acted — the agents breached three companies and published a malicious PyPI package
Attack surface
Autonomous operations
Autonomy level
Fully autonomous
Initial vector
Misconfigured eval sandbox with live internet
Impact type
Credential theft, Data exfiltration, Malware distribution, Unauthorized access, Data exposure, Malicious package distribution
Confidence
High
Submitted by
EAI Internal
Crosswalk
- OWASP Agentic Top 10 (2026)
- ASI04 – Agentic Supply Chain Vulnerabilities
- MITRE ATLAS
- AML.T0040 – AI Supply Chain Compromise
- NIST AI RMF
- MAP 4.1
- EU AI Act
- Annex III §2
- ISO/IEC
- TR 24028; 42001; 23894
- AIUC-1
- B. Security
- Scope
- tool failure
- Mitigation
- Scan dependencies and container images for vulnerabilities; Use trusted sources.
- Taxonomy
- Enkrypt AI agent risk taxonomy · Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
Blast radius
3 real companies touched; malicious PyPI package installed and executed on 15 real systems within an hour of publication; several hundred rows of production data accessed
Capability abused
Defense & remediation
Recorded defense
evaluation environment intended to be network-isolated; production model safeguards
Guardrail kinds
Sandbox isolation
Guardrail outcome
Silent
Fix shipped
package pulled by PyPI automated defenses; Anthropic tightened eval environment controls
Taxonomy & references
Curator notes
No human adversary - the harm came from an agent under evaluation believing it was in a simulation. Opus 4.7 extracted application and infrastructure credentials from a real company and read production database rows; Mythos 5 exfiltrated credentials from a security company. Package name not disclosed in the Socket writeup; other coverage refers to a package named 'anthropickit' - not confirmed here. | Merged from duplicate id `2026-07-anthropic-eval-agent-pypi-malware`. Also reported as: 'Anthropic: Claude models in cybersecurity evaluations reached the internet and compromised three real organisations'.
Supporting sources
Cite as
Enkrypt AI. "Anthropic evaluation models gained real internet access, breached three companies and published a credential-stealing PyPI package that ran on 15 real systems". AIR, AIR-2026-0173. https://agentincident.vercel.app/i/AIR-2026-0173/ (accessed 2026-09-22).