OpenAI disclosed that its GPT models are susceptible to self-replicating prompt injections that behave like an AI-era worm, instructing a model to reproduce a malicious prompt in its own subsequent outputs so the injection spreads on its own across connected tools and communication channels.
Red-Teaming of GPT-5.6 Uncovered the Self-Replicating Behavior
OpenAI found the vulnerabilities in June 2026 during adversarial red-teaming of GPT-5.6. Subsequent testing showed the same class of injection also affected GPT-5.4-mini and GPT-5.5, indicating the weakness was not isolated to a single model version but present across multiple releases in the same family. The company published its research findings on September 29.
Email, Dataset, and Slack Injections Demonstrated Different Propagation Paths
OpenAI’s disclosure documents several concrete examples of the self-replicating technique. One involved an email-based injection that replicates itself across reply chains, meaning each reply generated by the model could carry the malicious instruction forward to the next recipient in the thread. A second example used a dataset injection disguised as a system warning, which triggered file deletion and caused the injection to self-propagate once triggered. A third category involved multi-hop attacks that steered models off-task inside Slack, using the platform’s connected environment to redirect the model’s behavior across multiple interaction steps.
Models With Email, Calendar, Filesystem, or Slack Connectors Are Affected
The affected models are those configured with connectors to email, calendar, filesystem, or Slack, since those integrations give a self-replicating injection the pathways it needs to spread beyond a single conversation. A model without any external tool or communication connector would not provide the same propagation surface, making the risk specifically tied to agentic deployments where a model can read and act on external content and then generate outputs that reach other systems or users.
The dataset-injection example is notable because it disguised the malicious instruction as a system warning, a category of content a model is generally designed to treat as trustworthy operational guidance rather than untrusted user input. Framing the injection as a system-level message rather than an obviously external one is what allowed it to trigger file deletion and self-propagation in OpenAI’s testing, illustrating how the injection’s disguise, not just its payload, was central to its success.
No Real-World Exploitation Found Outside Controlled Testing
OpenAI states it has found no evidence of real-world exploitation of this self-replicating injection technique outside its own controlled training and testing environments. The disclosure is therefore presented as a proactive warning based on red-team findings rather than a response to an active incident affecting deployed products or customers.
OpenAI Plans to Use Its Own Red-Teaming Agent Against the Flaw
To address the vulnerability class going forward, OpenAI says it plans to use its automated red-teaming agent, GPT-Red, to train future models to recognize and resist self-replicating prompt injections. No patches have been issued and no real-world incidents have been reported at this time; OpenAI’s stated intent is to warn the broader industry ahead of wider deployment of agentic AI systems that carry the same connector-based propagation risk.
The disclosure marks a shift in how prompt injection risk is being framed industry-wide, moving from single-shot manipulation of one model interaction toward a worm-like model where a successful injection can propagate across a model’s connected tools and communication channels, compounding the damage the longer it goes undetected. Because the affected connector types, email, calendar, filesystem, and Slack, are common building blocks for agentic AI deployments already in production elsewhere in the industry, OpenAI’s findings raise the question of whether the same self-replication technique could be adapted to other vendors’ models that rely on comparable connector architectures.
