Skip to content
← All termsGlossary

Guardrails

Guardrails are the limits you put around an AI agent so it can work on its own without causing harm. They block risky actions and cap what the agent can reach or spend. Good guardrails let an agent do real work while keeping its blast radius small.

Also called: Agent guardrails, AI guardrails, Safety rails, Policy guardrails

Updated

An AI agent that can act on its own can also act wrong on its own. Guardrails are how you give it room to work without giving it room to do damage.

Why it matters

The value of an agent is that it acts without a human clicking approve every time. Guardrails make that safe. They set the outer edge, so a mistake stays small instead of turning into a breach.

Where teams get it wrong

Many teams treat a system prompt as a guardrail. A prompt is a request, not a limit. Real guardrails sit outside the model, in code the agent can't talk its way past. If a user or a poisoned page can argue the agent out of the rule, it was never a guardrail.

One example

Give a support agent a hard cap on refund size. Ask it nicely for more and it should still be blocked, because the cap lives in the payment system, not in the prompt. That's a guardrail doing its job.

Let's figure out what you actually need.

No pitch. No pressure. Every conversation starts with a senior practitioner, not a sales team. We'll tell you straight where you stand and whether we're the right fit. If we're not, we'll point you to who is.