Guardrails
Guardrails are the rules and checks placed around an AI agent to keep its behavior inside agreed limits: what it may discuss, what it must never say or do, and when it must hand off to a person.
Also called: safety rails, behavioral constraints, policy enforcement
A system prompt tells the model how to behave; guardrails make sure it does. They range from instructions ("never quote a price you have not been given") to hard mechanisms outside the model: filters on what it can say, limits on which actions it can take, checks on the input for manipulation attempts, and rules that force an escalation to a human when a topic is out of bounds, a medical emergency, a legal threat, an abusive caller.
The distinction matters because a prompt can be argued with and a mechanism cannot. Well-built agents layer the two: the prompt for tone and judgment, the mechanism for the lines that must not be crossed.
For a buyer, guardrails are the answer to "what stops it saying something we would never say?" A good product can list them.
