Indirect Prompt Injection: The Hidden Attack Vector in RAG & Agents 2026 AI Safety Directory

indirect prompt injection

In a RAG pipeline, user queries trigger retrieval of relevant documents from a knowledge base, and these documents are inserted into the model’s context window alongside the query. Attribution is harder for indirect injection because the malicious content arrives through a trusted data channel. The attack vector is the user input channel, and defenses focus on scanning and filtering user inputs before they reach the model. 1 Performance metrics are based on results from internal benchmark testing. Indirect prompt injection attacks are accessible to nation-states and individuals alike. Prompt injection attacks can enable adversaries to exfiltrate sensitive data, manipulate business processes, and conduct reconnaissance.

Even if indirect injection manipulates the model’s intent, human review can catch unauthorized actions before they execute. Indirect injection is considered higher severity because it bypasses input defenses, can be scaled (one poisoned document can affect many users), and is harder to detect in real time. This makes indirect injection fundamentally harder to defend against than direct injection, where the attacker’s input can be scanned and filtered in real time. Defending against indirect prompt injection attacks demands a multi-layered approach that addresses both technical controls and organizational processes to limit the size of the attack surface and detect and stop injection attacks. A recent New York Times article reported on a job applicant who manipulated an AI hiring platform with an indirect prompt injection attack and who “wrote more than 120 lines of code to influence A.I. And yet, this is not unlike what happens on a daily basis at organizations where both approved and unknown AI tools continually and indiscriminately crawl the web and internal resources, ingesting text, files, and multimedia assets that could contain indirect prompt injections.

While not a robust defense on its own (attackers can include fake delimiters in their payloads), it provides additional signal that helps the model distinguish instruction sources. Context tagging marks different parts of the prompt with clear delimiters that distinguish system instructions, user input, and retrieved data. A separate response model uses only the extracted information (facts, quotes, data points) without seeing the raw retrieved documents. In this design, a retrieval model extracts relevant information from documents but is restricted from generating final responses. Researchers demonstrated that a malicious Google Doc shared with a user could manipulate Bard’s responses when the user asked Bard questions about the document’s content.

  • End users of the AI tools targeted by indirect prompt injection will likely never see the malicious prompt, and the AI tool may even appear to function normally while subtly executing the attacker’s hidden instructions in the background.
  • When an LLM is used to evaluate the candidate, the combined prompts manipulate the model’s response, resulting in a positive recommendation despite the actual resume contents.
  • Enforce strict context adherence, limit responses to specific tasks or topics, and instruct the model to ignore attempts to modify core instructions.
  • Meta’s AI research division publishing open-source safety tools including LlamaGuard and LlamaFirewall.

Gemini AI

indirect prompt injection

This guide gets updated when the threat landscape shifts. Prompt injection is one of six questions worth asking of any AI system, set out in our AI security field guide. What that resilience requirement means in engineering terms, rather than legal terms, is covered in that guide. The EU AI Act requires high-risk AI systems to https://callmeconstruction.com/news/key-strategies-for-ctos-to-leverage-mern-stack-development-effectively/ be resilient against attempts to alter their intended purpose through manipulation of inputs.

Training Data Hygiene

Improper Output Handling refers specifically to insufficient validation, sanitization, and handling of the outputs generated by large language models… LLM supply chains are susceptible to various vulnerabilities, which can affect the integrity of training data, models, and https://financeswizards.com/revolutionize-business-methods.html deployment… When an LLM is used to evaluate the candidate, the combined prompts manipulate the model’s response, resulting in a positive recommendation despite the actual resume contents. A company includes an instruction in a job description to identify AI-generated applications.

indirect prompt injection

Direct Injection Examples

In February 2025, Ars Technica reported vulnerabilities in Google’s Gemini AI to indirect prompt injection attacks that manipulated its long-term memory. In December 2024, The Guardian reported that OpenAI’s ChatGPT search tool was vulnerable to indirect prompt injection attacks, allowing hidden webpage content to manipulate its responses. Threats to RAG systems including knowledge base poisoning, indirect prompt injection, and data exfiltration, with practical defenses. Data poisoning requires access to the training pipeline; indirect injection only requires the ability to place content where the model will retrieve it.

  • The EU AI Act requires high-risk AI systems to be resilient against attempts to alter their intended purpose through manipulation of inputs.
  • Indirect prompt injection is not theoretical — it has been demonstrated against production systems and extensively studied by security researchers.
  • Data poisoning modifies training data to permanently alter model behavior — the attack is embedded in the model’s weights.
  • Security researcher Johann Rehberger demonstrated how hidden instructions within documents could be stored and later triggered by user interactions.
  • While intentional and direct injection represents a threat to the developer from the user, unintentional indirect injection represents a threat from the data-author to the user.

Perform regular penetration testing and breach simulations, treating the model as an untrusted user to test the effectiveness of trust boundaries and access controls. Enforce strict context adherence, limit responses to specific tasks or topics, and instruct the model to ignore https://www.e-lib.info/getting-to-the-point-7/ attempts to modify core instructions. Provide specific instructions about the model’s role, capabilities, and limitations within the system prompt. Multimodal models may also be susceptible to novel cross-modal attacks that are difficult to detect and mitigate with current techniques.

Leave a Reply

Your email address will not be published. Required fields are marked *