Content guards & egress
Validating a tool call bounds what it can do; content guards bound what flows through it — the arguments going in and the results coming back. They ride the same pre- and post-tool-call pipeline as every other check, and they fail closed.
Pluggable inspectors on tool I/O
A content guard is an inspector that runs on a tool's input (before it runs) or its output (after), and returns one of three verdicts: pass, block, or raise a human approval. Inspectors are pluggable and self-registering, and the framework fails closed — an inspector that errors denies the call rather than waving it through — with a cap on how much content it will inspect so a payload can't be padded to outrun the detector. Two ship today: a prompt-injection detector and a URL allowlist.
Prompt injection, checked in a zero-egress pod
The prompt-injection detector runs a classifier over inspected text and, above a threshold, blocks the content or raises an approval. Because the text it's judging may itself be hostile, the classifier runs in its own per-session pod with egress denied — and that zero-egress policy is applied unconditionally, even where per-session network policies are otherwise turned off, because "the thing inspecting hostile text must not be able to phone home" is the whole point of the feature.
URL allowlists
The URL allowlist inspector extracts URLs from tool input and output and evaluates them against ordered rules, default-deny, strictest-match wins. It's how you keep a tool from fetching — or a result from smuggling in a link to — somewhere it shouldn't, before that URL ever reaches the model.
Beyond content: who may the data reach?
Content guards inspect what a tool's I/O contains. The related question — who is allowed to receive the data a tool read, once the agent goes to send it — is an authorization question, and it lives with the rest of them: see Information leakage & egress.