Quick Answer: LLM guardrails are checks that run before a large language model (input), after it (output) and around its actions (tool calls). Input and output guardrails catch harmful, leaked or malformed text, while tool-call controls and human approval limit what an agent can do. OWASP ranks prompt injection first in its 2025 LLM Top 10, and no single layer stops it.
A chatbot that only answers questions needs two checkpoints. An agent that can issue refunds or edit records needs a third, because a bad action costs more than a bad sentence.
This guide is for engineering and platform leads deciding which layer owns which risk, where each check runs, and which actions a person must still approve.
What are LLM guardrails?
LLM guardrails are programmable checks around a language model that block, rewrite or validate what goes in, what comes out and what the model may do. They run outside the model, so a fooled model cannot switch them off.
Production systems combine three layers rather than picking one:
- Input guardrails decide whether a request should reach the model at all.
- Output guardrails decide whether a response is safe to release.
- Tool-call controls decide whether the model may take an action, such as calling an API or writing to a database.
The OWASP Top 10 for LLM Applications 2025 maps onto the same split: prompt injection (LLM01) is mostly an input problem, improper output handling (LLM05) an output problem, and excessive agency (LLM06) a tool-call problem.
How do input, output and tool-call guardrails compare?
They differ in where they run and what they can stop. Input guardrails screen the request, output guardrails screen the text that comes back, and tool-call controls gate the action itself. Only the third layer can stop a harmful action that arrives through clean-looking text.
| Layer | What it catches | Example techniques | Latency cost | Failure mode | Who owns it |
|---|---|---|---|---|---|
| Input | Jailbreaks, prompt injection, personal data, off-policy topics | Pattern rules, classifiers, PII masking | Low; one extra model call for a classifier | Misses injection hidden in documents | Platform team |
| Output | Leaked secrets and PII, toxic text, broken JSON, unsupported claims | Schema validation, secret scanning, grounding checks | Low to moderate; can hold back streaming | Cannot undo an action already taken | Application team |
| Tool call | Unauthorized or destructive actions, bad arguments, runaway loops | Allowlists, scoped credentials, argument checks, dry runs | Near zero; longer when a person approves | Over-broad permissions; rubber-stamp approvals | Platform team plus system owners |
Where checks run matters too. Many teams enforce input and output rules once, at an LLM gateway that every internal app and coding assistant already calls, so one set of PII and injection rules covers every model. Tool-call controls usually live closer to the tools, in the agent runtime or the downstream API.
What do input guardrails catch before a prompt reaches the model?
Input guardrails catch requests that should never reach the model: jailbreak attempts, prompt injection, personal data the model should not see, and topics outside the application's policy.
- Injection and jailbreak detection: heuristics for known patterns, plus a small classifier for new phrasings.
- PII redaction: mask names, account numbers and health details before the prompt leaves your boundary.
- Topic filters: keep a support assistant on support.
- Untrusted content screening: check retrieved documents, emails and web pages, which can carry hidden instructions.
NVIDIA NeMo Guardrails, an open-source Python library, documents rails for each of these, including jailbreak detection, topic control and PII detection that can block or mask entities.
The common mistake is treating input filtering as a complete defense against prompt injection; OWASP says it is unclear whether any fool-proof prevention exists. Output filters cannot close the gap either, because they react after generation and cannot recall an API call that already ran.
What do output guardrails check before a response ships?
Output guardrails check that a response is safe, correctly shaped and free of data it should not reveal before anyone receives it.
- Schema validation: parse JSON against a schema and retry or reject on failure.
- Secret and PII scanning: catch API keys, connection strings and customer records pulled in through retrieval.
- Moderation: flag toxic or off-policy text.
- Grounding checks: compare claims with the retrieved sources and flag anything unsupported.
Guardrails AI, a Python framework, packages these checks as validators that you combine into input and output guards. OWASP's advice for LLM05 is to treat the model like any untrusted user and validate its output before it reaches backend functions. Grading response quality at scale is a related job, and how LLM judges compare with human reviewers covers it.
How do tool-call guardrails limit what an agent can do?
Tool-call guardrails sit between the model and the systems it can touch. They decide whether a proposed action runs, with which arguments and under whose credentials.
- Allowlists. Give the agent only the tools the task needs.
- Scoped credentials. The agent acts as the requesting user with the smallest scope, never through a shared admin key. NeMo's guidance is to isolate all authentication information from the LLM.
- Argument validation. Check IDs, amounts and recipients against rules written in code, not in the prompt.
- Rate limits. Cap calls per session so a loop cannot run up cost or flood a downstream API.
- Dry-run mode. Show the plan or diff first and commit only after it passes.
OWASP traces excessive agency to excessive functionality, permissions and autonomy; allowlists, scoped credentials and dry runs address each in turn. An agent needs this layer because it picks its own next action, unlike an RPA bot, as agentic AI versus RPA explains.
When should a human approve an agent's action?
A human should approve any action that is hard to reverse, moves money or data outside its normal path, or affects a person's health, finances or legal rights.
In regulated industries such as fintech and healthcare, a human-in-the-loop design means sorting every tool the agent can call into risk tiers before launch:
- Tier 1, read-only or easy to undo (look up an order): run automatically and log it.
- Tier 2, reversible writes within limits (update a ticket, refund under a set threshold): run automatically with sampled review.
- Tier 3, irreversible or regulated (move funds, change a medication order, deny a claim): pause for a named approver who sees the inputs, the proposed action and its sources.
NIST asks for this to be written down. In the NIST AI Risk Management Framework, MAP 3.5 reads: "Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function."
The Generative AI Profile (NIST AI 600-1) builds on GOVERN 3.2: "Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems." It also names automation bias, "excessive deference to automated systems", as a risk.
So log who approved what, when and on which inputs, and track approval rates: an approver who accepts nearly everything is no longer a control.
What mistakes should you avoid when layering LLM guardrails?
- Treating the system prompt as a guardrail. OWASP's LLM07 entry says a system prompt should not be used as a security control. Rules that matter belong in code.
- Relying on one layer. Input filters miss indirect injection, and output filters arrive too late for actions.
- Skipping the logs. Without a record of every block, pass and approval, you cannot tune false positives or show that oversight happened.
- Blocking so much that users route around it. If a filter rejects routine work, people move to personal accounts where nothing is checked.
- Testing only the happy path. Red-team with instructions hidden inside documents and tool results.
Design approval steps with the workflow, not after launch, as part of any AI workflow and agent engineering project.
How Origins AI layers guardrails in production
Origins AI (originshq.com) builds custom AI workflows and agents and deploys its self-hosted AI products inside the customer's environment. Two of its product pages describe controls for those layers.
The gateway layer comes from the Origins AI Coding Tool, a self-hosted LLM gateway and coding assistant. According to its product page, it redacts secrets, API keys and PII before content reaches the model, enforces per-team access policies and token quotas, and logs every request and response in your environment. The same page states that in on-premise and air-gapped modes no source code is sent to an external service, while hybrid mode sends only the submitted context to the model.
For agents that act, the Origins AI Agentic Automation page describes an agent-design step that sets autonomy boundaries, escalation triggers and approval thresholds before a pilot on one process.
The security controls Origins AI lists for its services are encryption at rest and in transit, secure authentication, continuous security monitoring and least-privilege access.
Talk to an engineer
Planning guardrails for an agent that will touch production systems? Book a call with an Origins AI engineer to map your agent's actions to risk tiers and decide where each check should run.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


