Safe tools
An LLM agent is only as safe as the tools you hand it. When the thing that runs a tool call is a Python function with full process authority, the tool surface is whatever code happens to expose, arguments go unvalidated, and every tool runs with ambient authority. OAP restricts an agent to sandboxed CLI tools and MCP servers, validated against a toolspec — it can only do what was written down.
How it works
A tool in OAP is a specification, not a function. Each is declared with a deny-by-default allowlist of
subcommands and fields (refined with CEL), and it is validated twice — once at authoring time, and again at
execution time, so a spec that drifts can't quietly widen. Tools run as argv arrays over a sandboxed pod —
never through a shell, never eval — so there's no string to inject into.
Every tool declares its stateImpact — readonly, readwrite, or external — and that drives both the
authorization check and whether a human is asked before it runs. The sandbox pods themselves are hardened
(non-root, read-only root filesystem, no ambient service-account token), and default-on circuit breakers,
plus optional rate limits and per-call data-volume budgets, keep a misbehaving tool from running away.
Go deeper
- Toolspec & validation — deny-by-default allowlists, CEL, and dual validation.
- Sandboxed execution — the hardened pod model and argv-not-shell execution.
- MCP servers & sidecar toolboxes — bringing external tools under the same rules.
- Content guards & egress — inspecting tool I/O and bounding what leaves the session.
- Budgets, rate limits & circuit breakers — bounding runaway use.
- Supply-chain pinning — content-hashing the tools an agent is allowed to run.