Contact Us

LLM Guardrails for Inputs, Outputs and Tool Calls (2026)

Sep 25, 20267 min read
Light beams pass through three glass gates to a server rack: LLM Guardrails for Inputs, Outputs and Tool Calls (2026)
llm guardrails input guardrails output guardrails prompt injection

TL;DR

  • Only tool-call controls can stop a harmful action that arrives through clean-looking text, because output filters cannot recall an API call already made.
  • A system prompt should not be used as a security control, so the rules that matter belong in code outside the model.
  • Require a named approver for irreversible or regulated actions such as moving funds, and track approval rates to catch rubber-stamp approvals.

Quick Answer: LLM guardrails are checks that run before a large language model (input), after it (output) and around its actions (tool calls). Input and output guardrails catch harmful, leaked or malformed text, while tool-call controls and human approval limit what an agent can do. OWASP ranks prompt injection first in its 2025 LLM Top 10, and no single layer stops it.

A chatbot that only answers questions needs two checkpoints. An agent that can issue refunds or edit records needs a third, because a bad action costs more than a bad sentence.

This guide is for engineering and platform leads deciding which layer owns which risk, where each check runs, and which actions a person must still approve.

What are LLM guardrails?

LLM guardrails are programmable checks around a language model that block, rewrite or validate what goes in, what comes out and what the model may do. They run outside the model, so a fooled model cannot switch them off.

Production systems combine three layers rather than picking one:

The OWASP Top 10 for LLM Applications 2025 maps onto the same split: prompt injection (LLM01) is mostly an input problem, improper output handling (LLM05) an output problem, and excessive agency (LLM06) a tool-call problem.

How do input, output and tool-call guardrails compare?

They differ in where they run and what they can stop. Input guardrails screen the request, output guardrails screen the text that comes back, and tool-call controls gate the action itself. Only the third layer can stop a harmful action that arrives through clean-looking text.

Layer What it catches Example techniques Latency cost Failure mode Who owns it
Input Jailbreaks, prompt injection, personal data, off-policy topics Pattern rules, classifiers, PII masking Low; one extra model call for a classifier Misses injection hidden in documents Platform team
Output Leaked secrets and PII, toxic text, broken JSON, unsupported claims Schema validation, secret scanning, grounding checks Low to moderate; can hold back streaming Cannot undo an action already taken Application team
Tool call Unauthorized or destructive actions, bad arguments, runaway loops Allowlists, scoped credentials, argument checks, dry runs Near zero; longer when a person approves Over-broad permissions; rubber-stamp approvals Platform team plus system owners

Where checks run matters too. Many teams enforce input and output rules once, at an LLM gateway that every internal app and coding assistant already calls, so one set of PII and injection rules covers every model. Tool-call controls usually live closer to the tools, in the agent runtime or the downstream API.

What do input guardrails catch before a prompt reaches the model?

Input guardrails catch requests that should never reach the model: jailbreak attempts, prompt injection, personal data the model should not see, and topics outside the application's policy.

NVIDIA NeMo Guardrails, an open-source Python library, documents rails for each of these, including jailbreak detection, topic control and PII detection that can block or mask entities.

The common mistake is treating input filtering as a complete defense against prompt injection; OWASP says it is unclear whether any fool-proof prevention exists. Output filters cannot close the gap either, because they react after generation and cannot recall an API call that already ran.

What do output guardrails check before a response ships?

Output guardrails check that a response is safe, correctly shaped and free of data it should not reveal before anyone receives it.

Guardrails AI, a Python framework, packages these checks as validators that you combine into input and output guards. OWASP's advice for LLM05 is to treat the model like any untrusted user and validate its output before it reaches backend functions. Grading response quality at scale is a related job, and how LLM judges compare with human reviewers covers it.

How do tool-call guardrails limit what an agent can do?

Tool-call guardrails sit between the model and the systems it can touch. They decide whether a proposed action runs, with which arguments and under whose credentials.

  1. Allowlists. Give the agent only the tools the task needs.
  2. Scoped credentials. The agent acts as the requesting user with the smallest scope, never through a shared admin key. NeMo's guidance is to isolate all authentication information from the LLM.
  3. Argument validation. Check IDs, amounts and recipients against rules written in code, not in the prompt.
  4. Rate limits. Cap calls per session so a loop cannot run up cost or flood a downstream API.
  5. Dry-run mode. Show the plan or diff first and commit only after it passes.

OWASP traces excessive agency to excessive functionality, permissions and autonomy; allowlists, scoped credentials and dry runs address each in turn. An agent needs this layer because it picks its own next action, unlike an RPA bot, as agentic AI versus RPA explains.

When should a human approve an agent's action?

A human should approve any action that is hard to reverse, moves money or data outside its normal path, or affects a person's health, finances or legal rights.

In regulated industries such as fintech and healthcare, a human-in-the-loop design means sorting every tool the agent can call into risk tiers before launch:

NIST asks for this to be written down. In the NIST AI Risk Management Framework, MAP 3.5 reads: "Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function."

The Generative AI Profile (NIST AI 600-1) builds on GOVERN 3.2: "Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems." It also names automation bias, "excessive deference to automated systems", as a risk.

So log who approved what, when and on which inputs, and track approval rates: an approver who accepts nearly everything is no longer a control.

What mistakes should you avoid when layering LLM guardrails?

Design approval steps with the workflow, not after launch, as part of any AI workflow and agent engineering project.

How Origins AI layers guardrails in production

Origins AI (originshq.com) builds custom AI workflows and agents and deploys its self-hosted AI products inside the customer's environment. Two of its product pages describe controls for those layers.

The gateway layer comes from the Origins AI Coding Tool, a self-hosted LLM gateway and coding assistant. According to its product page, it redacts secrets, API keys and PII before content reaches the model, enforces per-team access policies and token quotas, and logs every request and response in your environment. The same page states that in on-premise and air-gapped modes no source code is sent to an external service, while hybrid mode sends only the submitted context to the model.

For agents that act, the Origins AI Agentic Automation page describes an agent-design step that sets autonomy boundaries, escalation triggers and approval thresholds before a pilot on one process.

The security controls Origins AI lists for its services are encryption at rest and in transit, secure authentication, continuous security monitoring and least-privilege access.

Talk to an engineer

Planning guardrails for an agent that will touch production systems? Book a call with an Origins AI engineer to map your agent's actions to risk tiers and decide where each check should run.

Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.

Frequently Asked Questions

What do real-world AI guardrail examples look like?
Typical ones are a jailbreak classifier on incoming prompts, PII masking before a prompt leaves your network, a topic filter on a support bot, JSON schema validation on responses, secret scanning on outputs, a short allowlist of tools an agent may call, and a human sign-off step before any payment or record deletion.
Does ChatGPT have guardrails?
Yes, at the provider level. For API developers, OpenAI documents moderation models that detect harmful content in text and images, with results used to filter content or route a request for review. Those checks do not know your data rules, your tools or which actions need sign-off, so your application still needs its own layers.
What are guardrails for AI agents?
They are the controls between an agent and the systems it can change: a short list of permitted tools, credentials scoped to the requesting user, argument checks in code, call limits per session, and human approval for irreversible steps. Text filters alone cannot stop an action.
What is prompt injection?
Prompt injection is an attack where text in the model's input changes its behavior in ways the developer did not intend. Direct injection comes from a user typing instructions. Indirect injection hides instructions inside content the model reads, such as a web page, an email or a retrieved document, so it often slips past filters that watch only the chat box.
Can guardrails stop prompt injection completely?
No. The weakness comes from how models process text, and OWASP says no fool-proof prevention is known. Guardrails lower the rate and limit the damage. Aim for containment: filter inputs, validate outputs, give the agent the least privilege it needs, and require a person to approve high-impact actions.
Do guardrails slow down LLM responses?
Some do. Rule-based checks add almost nothing, a classifier or second model call adds a step per request, and output checks that need the full answer can delay streaming. Measure p95 latency with real prompts, run cheap checks first, and keep model-based checks for risky requests.
Book a call

About the Author

Apoorva Kumar is Co-Founder and CEO of Origins AI (originshq.com), an AI engineering partner for product teams building AI workflows, AI agents and LLM integrations. A CSE graduate of IIT Kharagpur, Apoorva previously built and scaled technology at Sony, NuCash, YesMadam and FrontPage.