Vulnerability  ·  2026-09-27

OpenAI discloses self-replicating prompt injections ('AI worm') found via GPT-Red — first lab confirmation of worm-class agentic prompt injection (disclosed 2026-09-25)

VulnerabilityMedium impactGlobal
On 2026-09-25 OpenAI's Alignment team published 'Self-replicating prompt injections exist' (discovery 2026-06-27) demonstrating via the GPT-Red self-play framework that prompt injections can self-propagate akin to a computer worm: via email replay, filesystem writes, and code comments, using fake chain-of-thought and fake-tool-message styles plus multi-hop chains. OpenAI states no impact occurred outside simulated environments and released the finding due to the novel nature of the attack, not an incident.
This is the first major AI lab publicly acknowledging worm-class self-replicating prompt injection in its own models — a novel agent-execution attack class with a demonstrated, reproducible mechanism. It validates that agent connectors (email/calendar/files/code-repos) are a self-propagating channel, the exact architecture enterprises are deploying today. Defenders should treat all ingested context (email, files, comments, summaries) as untrusted instructions and enforce separate-instruction handling at the harness/trust layer.
A prompt injection that both achieves an adversarial goal and induces the agent to reproduce the injected instructions on a public/sharable output channel (e.g. append a literal quote of the entire email body to every outgoing reply). Copying the payload into output lets it propagate to any downstream agent that reads that email/file/comment — a worm-replication primitive via connectors like email and calendar.
No product patch; OpenAI is using GPT-Red to harden models. Enterprise mitigation: keep untrusted data separate from instructions (system-prompt isolation), independent authorization per tool call, outbound content filtering, and treat any agent output that embeds quotations of ingested content as a potential propagation signal.
OpenAI Alignment - Self-replicating prompt injections existShattered.io analysis of OpenAI disclosureCryptoBriefing coverage
See this in the live feed Explore related AI security and governance findings — updated every morning.
Open the feed →