Safe tools

An LLM agent is only as safe as the tools you hand it. When the thing that runs a tool call is a Python function with full process authority, the tool surface is whatever code happens to expose, arguments go unvalidated, and every tool runs with ambient authority. OAP restricts an agent to sandboxed CLI tools and MCP servers, validated against a toolspec — it can only do what was written down.

Read-only steps run freely; the external deploy asks for approval first, then runs in a locked-down sandbox.

How it works

A tool in OAP is a specification, not a function. Each is declared with a deny-by-default allowlist of subcommands and fields (refined with CEL), and it is validated twice — once at authoring time, and again at execution time, so a spec that drifts can't quietly widen. Tools run as argv arrays over a sandboxed pod — never through a shell, never eval — so there's no string to inject into.

Every tool declares its stateImpact — readonly, readwrite, or external — and that drives both the authorization check and whether a human is asked before it runs. The sandbox pods themselves are hardened (non-root, read-only root filesystem, no ambient service-account token), and default-on circuit breakers, plus optional rate limits and per-call data-volume budgets, keep a misbehaving tool from running away.

Go deeper