Quick Answer: You can't fully prevent prompt injection in AI agents, so you prevent the damage in the agent's architecture rather than its prompt. Four controls carry most of the load: least-privilege tools, untrusted content kept apart from instructions, validated outputs and human approval before risky actions. OWASP lists prompt injection as LLM01, the top risk in its 2025 LLM Top 10.
If you're working out how to prevent prompt injection before an agent goes live, start from one fact: the model reads instructions and data through the same channel, and no filter separates them reliably. NIST's 2025 adversarial machine learning taxonomy says current mitigations "do not offer full protection" and that designers may build on the assumption that injection is possible whenever a model reads untrusted input.
What is prompt injection and why do agents make it worse?
Prompt injection is input that changes what a language model does in ways its builder didn't intend, because the model can't reliably tell trusted instructions from the text it's processing. OWASP's LLM01 entry adds that the injected text doesn't need to be visible to a person, only parsed by the model.
In a chatbot, prompt injection attacks mostly produce bad text. In an agent, the same manipulation becomes an action. OWASP lists the outcomes once tools are attached: unauthorized access to functions the model can call, commands executed in connected systems and manipulated decisions.
So severity depends on the agent's permissions. An agent that can send email, move money or merge code can do those things for an attacker.
What is the difference between direct and indirect prompt injection?
Direct injection comes from the person typing to the model. Indirect prompt injection arrives inside content the agent fetches, such as an email, a web page or a tool result.
- Direct: the user is the attacker, typing instructions meant to override the system prompt.
- Indirect: a third party plants instructions in something the agent will read. Microsoft's MSRC write-up lists hidden text on web pages, emails, shared documents and data returned by tools, and notes that even a plain .txt file can carry it.
Indirect injection is harder to catch because, as NIST notes, the primary user is often the victim, not the attacker. Microsoft reports it's one of the most widely used techniques in AI vulnerabilities submitted to it.
Which defenses reduce prompt injection risk in AI agents?
No single control holds, so you stack layers that each assume the one before it failed. The table follows OWASP's guidance and its prompt injection prevention cheat sheet.
| Layer | What it stops | Example control | Residual risk |
|---|---|---|---|
| Input handling | Known injection phrasing and encoded variants | Pattern and fuzzy-match filters, an input classifier | Rephrased or novel attacks |
| Content isolation | Retrieved text read as instructions | Labeled channels for untrusted content; delimiting, datamarking or encoding | Models still sometimes follow marked text |
| Least-privilege tools | Injected instructions reaching sensitive systems | Read-only credentials, per-tool scopes, allow-listed actions | Damage inside the granted scope |
| Output validation | Data exfiltration and malformed tool calls | Schema checks in code, no rendering of external images or links | Valid-looking but wrong output |
| Human approval | Irreversible actions on injected instructions | An approval gate on payments, deletions and outbound messages | Approval fatigue |
| Monitoring and logging | Slow, repeated or silent attacks | Logs of every prompt, retrieval and tool call, with alerts | Detection comes after the fact |
Favor deterministic controls: a parameter check in code blocks outright, while a filter only lowers the odds. Model-based LLM guardrails can themselves be injected, so count them as one layer.
Routing agent traffic through one LLM gateway gives you a single place to log prompts, filter secrets and enforce model access.
How do you limit what an agent can do if it is injected?
You limit an injected agent by giving it the smallest set of tools and credentials the task needs, and by putting a human decision in front of any action that's costly or hard to undo.
Scope credentials and tools
- Give the agent its own service account, never an employee's token or an admin key. OWASP advises handling privileged functions in code rather than handing them to the model.
- Default to read-only, grant write scopes per tool and allow-list recipients, domains and repositories.
- Separate reading from acting: in the dual-LLM pattern, a quarantined model reads untrusted content but can't act, and a privileged model acts but never sees that raw text.
Agents that reach tools through MCP servers add their own permission questions, covered in our comparison of MCP vs API security.
Add human approval where the stakes are regulated
In regulated workflows like fintech and healthcare, the strongest control is a human approval step before any consequential action: a payment, a credit decision, a change to a patient record or a message to a customer. Four design rules make that step count:
- Tier actions by risk. Automate low-risk reads, review exceptions at medium risk and require a qualified person's decision at high risk.
- Show the evidence. The reviewer sees the request, the sources the agent read and the exact action.
- Log the decision. Record who approved, what they saw and why they overrode anything.
- Name who can stop the agent. The NIST AI RMF core asks organizations to document human oversight processes (Map 3.5) and assign responsibility to disengage systems that behave outside intended use (Manage 2.4).
How do you test an agent for prompt injection?
You test an agent by attacking it through every channel it reads, on every change.
- Cover each attack class: direct, indirect, encoded, multi-turn, retrieval poisoning and manipulated tool output.
- Plant injected documents in a test copy of each source the agent reads, such as an inbox or a ticket queue.
- Use canaries. If a planted fake secret or harmless marker instruction shows up in output or outbound traffic, a boundary failed.
- Assert on actions. A test fails when the agent calls a tool it shouldn't, not only when it says something odd.
- Try many variants. Research cited in the OWASP cheat sheet reached 89% attack success on GPT-4o with enough rephrased attempts.
In production, alert on unusual tool calls.
What should a prompt injection review cover before launch?
A pre-launch review should prove you know every tool, every path untrusted text takes, where people approve, what's logged and how you'd shut the agent down.
- Tool inventory: each tool, its credential, its scope and the worst action it allows.
- Untrusted input map: every source the agent reads and how each is isolated.
- Approval gates: which actions need a person, who approves and what they see.
- Output controls: link and image rendering rules, schema validation and an egress allow-list.
- Logging: prompts, retrieved documents, tool calls and approvals.
- Test evidence: the injection suite's last run date, pass rate and open failures.
- Incident plan: how to revoke the agent's credentials, disable its tools and roll back its actions.
If the agent plugs into existing enterprise systems, connect it through scoped service accounts and existing APIs, so you can revoke its access alone and roll it out one workflow at a time without disrupting operations.
What mistakes should you avoid when defending against prompt injection?
The most expensive mistake when you try to prevent prompt injection is treating the system prompt as a security boundary. Telling the model to ignore embedded instructions helps a little, but attackers iterate.
- Broad API keys. An agent holding an admin key turns one injection into a full compromise.
- Trusting your own index. Internal content is still untrusted if outsiders can write to its sources.
- Rendering output blindly. Auto-loaded images and links are a common exfiltration path.
- One-off testing. A suite that passed before launch says nothing after the next model change.
How Origins AI hardens agents against prompt injection
Origins AI (originshq.com) is an AI-augmented engineering company that builds agents and automated workflows with human-in-the-loop controls, and its product and services pages list the security controls below. None removes injection risk; they narrow what an injected agent reaches and record what it did.
The Origins AI Coding Tool includes a self-hosted LLM gateway that proxies model requests from engineering tools, with routing, rate limits and access control. Its product page says every request and response is logged in your environment, and that secrets and PII are redacted before content reaches the model.
It deploys on-premise, in your own cloud account, air-gapped or hybrid. In on-premise and air-gapped modes, no source code goes to an external service; hybrid mode sends the submitted code context to the hosted model.
For custom agents, the AI engineering services page lists AI agent deployment plus encryption at rest and in transit, secure authentication, continuous monitoring and least-privilege access. The Origins AI Agentic Automation page describes a design stage that sets autonomy boundaries, escalation triggers and approval thresholds before a pilot.
Talk to an engineer
Launching an agent that can take actions? Bring the pre-launch review checklist above and book a security review of its tools, data paths and approval steps.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


