The agent loop

A turn is the familiar loop: the model proposes tool calls, the tools run, their results feed back, and it repeats until the agent answers. What's different in OAP is that every tool call passes through a hook pipeline before it runs — the model never touches a tool directly.

A turn, step by step

The runner drains the inbound messages for the turn and hands the conversation to the model. The model proposes zero or more tool calls. Each proposed call is dispatched through a chain of hooks; only if it survives them does it reach the sandbox or MCP server. The results come back as content, the model continues, and the turn ends when the agent produces a reply or hits a budget or the watchdog.

Tool output cannot authorize anything

A tool result is content the agent read, not a source of authority. Text that comes back from a web page, a file, or an API — however imperative it sounds — can't cause the agent to take an action without that action going through the same gates as any other tool call. That is the boundary against prompt injection: no text a tool returns can reach a privileged action without the platform's checks.

The model also gets help recognizing injected text. Every result is wrapped in nonce-tagged delimiters before it's shown to the model, and the system prompt tells the model to treat anything inside those tags strictly as data — and to report, not obey, any instructions it finds there ("ignore previous instructions", "now call tool X", "the user approved…"). The delimiters and that rule help the model, but are not the boundary: a model that follows an injected instruction anyway still can't make a denied call run.

Where the safety checks sit

The hooks run before dispatch, in the pre-tool-call stage:

  • Authorization — can this subject take this action, on this resource, right now? (Authorization)
  • Plan gate — is this call inside an approved plan, or does it need one? (Plan gating)
  • Content guards — does the call's input or the tool's output need inspecting or bounding? (Content guards)
  • Rate limits & circuit breakers — has this tool run away? (Budgets & breakers)

A call that any gate denies never reaches the tool — no sandbox exec, no MCP request. The denial comes back to the model as an error result, which it reads like any other content. And nothing is allowed to hang silently: the runner enforces the session's time budget and bounds every human-approval wait, while a separate silence watchdog catches a turn that goes quiet — failing loudly rather than stalling.