What happened
On 2026-09-25 OpenAI's Alignment team published 'Self-replicating prompt injections exist' (discovery 2026-06-27) demonstrating via the GPT-Red self-play framework that prompt injections can self-propagate akin to a computer worm: via email replay, filesystem writes, and code comments, using fake chain-of-thought and fake-tool-message styles plus multi-hop chains. OpenAI states no impact occurred outside simulated environments and released the finding due to the novel nature of the attack, not an incident.
Why it matters
This is the first major AI lab publicly acknowledging worm-class self-replicating prompt injection in its own models — a novel agent-execution attack class with a demonstrated, reproducible mechanism. It validates that agent connectors (email/calendar/files/code-repos) are a self-propagating channel, the exact architecture enterprises are deploying today. Defenders should treat all ingested context (email, files, comments, summaries) as untrusted instructions and enforce separate-instruction handling at the harness/trust layer.
Attack vector
A prompt injection that both achieves an adversarial goal and induces the agent to reproduce the injected instructions on a public/sharable output channel (e.g. append a literal quote of the entire email body to every outgoing reply). Copying the payload into output lets it propagate to any downstream agent that reads that email/file/comment — a worm-replication primitive via connectors like email and calendar.
Mitigation
No product patch; OpenAI is using GPT-Red to harden models. Enterprise mitigation: keep untrusted data separate from instructions (system-prompt isolation), independent authorization per tool call, outbound content filtering, and treat any agent output that embeds quotations of ingested content as a potential propagation signal.