Self-replicating prompt injections exist · OpenAI Alignment

Self-replicating prompt injections exist · OpenAI Alignment

  • September 29, 2026
Table of Contents
Self-replicating prompt injections exist · OpenAI Alignment

Research on aligning AI with human values and intent, and reports documenting model failures. We show the existence of a new variety of prompt injection, which can self-propagate akin to a computer worm. No impact was observed outside of the simulated tool calls in training and evaluation; we are sharing this due to the novel nature of the prompt injection, not because of any incident.

Source: alignment.openai.com

Tags :
Share :
comments powered by Disqus

Related Posts

How LLMs Actually Work

How LLMs Actually Work

This post is a walkthrough of how LLMs work. Modern LLMs are mostly built by stacking transformer blocks over and over, so understanding the transformer machinery gets you most of the way there.I’ll cover the core mechanisms inside modern transformer-based LLMs, without all that sticky math stuff. Don’t get me wrong, you should learn the math, but this can serve as an introduction.Most modern LLMs share the same transformer-family skeleton. The differences come from what each one was trained on, the scale and configuration choices, and the post-training done on top. By the end, you should be able to read many modern LLM papers or model cards and know which piece of the architecture each section is talking about.

Read More
5 Essential Papers on AI Training Data

5 Essential Papers on AI Training Data

Many data scientists claim that around80% of their time is spent on data preprocessing, and for good reasons, as collecting, annotating, and formatting data are crucial tasks in machine learning. This article will help you understand the importance of these tasks, as well as learn methods and tips from other researchers. Below, we will highlight academic papers from reputable universities and research teams on various training data topics.

Read More