AgentSettings

Group agentprimitives.authzed.com · Scope Namespaced · Short names agset

AgentSettings is the namespace tier of agent governance settings: limits that narrow what agents in this namespace may do, and defaults that fill in fields a lower tier omits. Per-namespace singleton -- the only permitted name is "default" (AgentSettingsName).

Namespaced. Reconciled by pkg/controllers/settings, which reports SelfConsistent. Tier resolution itself lives in pkg/platform/settings: a limit here can only narrow the ClusterAgentSettings ceiling, never widen it, and the resolved snapshot is stamped onto AgentClass/AgentSession status.

Spec

FieldTypeDescription
spec.classUserPreferencesmap[string]map[string]objectClassUserPreferences carries admin globals for classes' userPreferences, keyed by class name (same namespace), then preference key. NAMESPACE TIER ONLY: the ClusterAgentSettings webhook rejects a non-empty value — namespaced class names make a cluster-tier global ambiguous.
spec.defaultsobjectDefaults are fallback values used when a lower tier omits the field.
spec.defaults.authzobjectAuthz defaults the per-class approval/authz timeouts.
spec.defaults.authz.approvalTimeoutstringApprovalTimeout defaults AuthzBlock.ApprovalTimeout.
spec.defaults.authz.informationLeakageApprovalTTLstringInformationLeakageApprovalTTL defaults InformationLeakagePolicy.ApprovalTTL.
spec.defaults.authz.metaagentobjectMetaagent defaults AuthzBlock.Metaagent — what a class gets when it says nothing. A class may still opt itself into inline; forbidding that would need a SettingsLimits floor, which does not exist for this yet.
spec.defaults.authz.metaagent.triggerstringTrigger is when the metaagent classifies a turn, and whether it acts. mention (default) — ambient off; explicit @metaagent only. shadow — classify every turn, apply nothing, record it. inline — classify every turn and act. (enum: mention | shadow | inline; default: mention)
spec.defaults.authz.planGateobjectPlanGate defaults AuthzBlock.PlanGate — what a class gets when it says nothing. A class may still declare a laxer mode; forbidding that is SettingsLimits.MinPlanGateMode's job, and is a separate deliberate act.
spec.defaults.authz.planGate.examples[]objectExamples are worked plans this class's author supplies, rendered after the one the runtime derives from the session's own permission surface. Optional, and most classes want none: the derived example already speaks the agent's own handles and shows the shape that costs one approval. These are for a class whose GOOD plan is not obvious from its surface alone — where the order of phases matters, or where two resource types must be named together, or where the natural split is by audience rather than by read-then-write. They are TRUSTED text, at the same level as spec.systemPrompt: an operator authors them, never the agent. But unlike the system prompt they are checked — every handle an example names must exist on this class's own surface, or the class goes Valid=False naming the offender. The prompt tells agents that a handle off the declarable list is silently dropped, and an EXAMPLE carrying such a handle would teach exactly that failure to every plan the agent writes. Measured, planners transcribe these examples closely, so a wrong one is not inert.
spec.defaults.authz.planGate.examples[].phases *[]objectPhases are the plan itself. Rendered as the JSON an agent would pass to update_plan, because that is the form it copies.
spec.defaults.authz.planGate.examples[].phases[].id *stringID is the phase's identifier, as update_plan takes it.
spec.defaults.authz.planGate.examples[].phases[].labelstringLabel is the human phrase for the phase.
spec.defaults.authz.planGate.examples[].phases[].permissions[]stringPermissions are the wire-format handles this phase declares ("perm:<permission>:<resourceType>" or "tool:<name>"). Each is checked against the class's surface at admission.
spec.defaults.authz.planGate.examples[].phases[].slots[]stringSlots are the resource TYPES this phase names an instance of. Types only: an authored example must not carry a concrete instance id, because a planner that copies one would address an object belonging to whoever the example was written about.
spec.defaults.authz.planGate.examples[].task *stringTask is the one-line request this plan answers — the "when you are asked to X" half. Without it an example shows a shape with no occasion, and the agent cannot tell which of several examples applies.
spec.defaults.authz.planGate.limitsobjectLimits bound plan size. The defaults are deliberately expansive: they exist to stop a runaway or hostile plan from DoS-ing the approver and producing an unrenderable card, NOT to shape normal authoring.
spec.defaults.authz.planGate.limits.maxPermissionsPerPhaseinteger (int32)MaxPermissionsPerPhase is naturally bounded by the permission surface (a phase cannot declare a handle that does not exist), so this is a backstop rather than a real constraint. (default: 64)
spec.defaults.authz.planGate.limits.maxPhasesinteger (int32)(default: 50)
spec.defaults.authz.planGate.limits.maxSlotsPerPhaseinteger (int32)MaxSlotsPerPhase bounds the instance axis, which is genuinely unbounded (a thread can carry many URLs), hence the larger default. (default: 128)
spec.defaults.authz.planGate.modestringMode selects how much of the gate runs. disabled — off entirely; no prompt change, no plan schema surfaced. logging — the gate runs in full: enumeration, ceiling computation, approver resolution, severity, card rendering and the append-only write. What it holds back is CEILINGS and APPROVALS — the card is logged instead of published, so nothing is approved and no grant is written, and a call that overruns a declared phase is RECORDED rather than refused. A call made with NO plan at all is still refused (see RequirePlan). That is not an exception to the mode so much as its precondition: without a declared plan there are no ceilings to record and the mode measures nothing. enforcing — the card is published and the call blocks on the outcome. (enum: disabled | logging | enforcing; default: disabled)
spec.defaults.authz.planGate.renderingobjectRendering bounds what a single approval card shows or auto-approves.
spec.defaults.authz.planGate.rendering.maxAutoApproveHandlesinteger (int32)MaxAutoApproveHandles bounds the UNION of all handles auto-approved without a human across a session — the tier-0 budget. (default: 8)
spec.defaults.authz.planGate.rendering.maxNamedApproversinteger (int32)MaxNamedApprovers is the largest approver population rendered by name before falling back to a count or a channel-supplied description. (default: 5)
spec.defaults.authz.planGate.rendering.maxSingleCardHandlesinteger (int32)MaxSingleCardHandles bounds a phase ceiling shown on one card; a phase exceeding it is split rather than folded behind "+N more", so the approver sees what they approve. (default: 16)
spec.defaults.authz.planGate.requirePlanbooleanRequirePlan denies every permissioned call made before a plan is declared. Without it, an agent that simply never calls update_plan falls through the implicit-phase rule and is unconstrained — i.e. "don't plan" is a total bypass. It bites in LOGGING mode too, which is the one place this differs from everything else the gate does. "You must declare a plan" and "you must stay within its ceiling" are separable claims; only the second is what logging mode holds back. Welding them together made logging unable to produce the dataset it exists for — agents told they MUST plan simply did not, because in logging nothing depended on it, and no wording fixes a consequence the code does not implement. So: with RequirePlan, a session under logging still refuses calls made with no plan at all, and still records-but-permits calls that overrun a declared ceiling. That yields real phases, ceilings, severity and approver data without switching on the enforcement the data is meant to justify. UNSET derives from Mode: required whenever the gate runs at all. Turning the gate on IS the intent to make agents plan, and leaving this an independent opt-in defaulting false meant a cluster could run mode: logging, look gated, and collect nothing — which is what happened across five live sessions, every call landing on the implicit phase because nothing required a plan. Explicit false remains an opt-out, because gate-on-but-observe-only is a real rollout stage: it records which handles calls actually resolve to without changing what the agent may do. Explicit true under mode: disabled is still false — a gate that does not run cannot refuse anything, and honouring it would promise a denial that never arrives. Pointer rather than a kubebuilder default: a static default cannot express "depends on another field", and it would erase the difference between unset and a deliberate false. UPGRADE NOTE. This field previously carried +kubebuilder:default=false, so every AgentClass created under the older CRD has an explicit false PERSISTED — which now reads as the observe-only opt-out rather than as unset. Such a class keeps the old behaviour after an upgrade and will not require a plan. Re-applying the class (or deleting the field) is what restores derivation; there is no way to distinguish a stamped default from a deliberate choice after the fact.
spec.defaults.authz.scopeMaxLlmLatencyMsinteger (int32)ScopeMaxLLMLatencyMs defaults ScopeSpec.MaxLLMLatencyMs.
spec.defaults.budgetobjectBudget defaults the agent budget when an AgentClass omits its own.
spec.defaults.budget.maxDelegatedAgentsinteger (int32)MaxDelegatedAgents bounds the TOTAL number of sessions in one delegation tree, counting the root. It is a property of the tree rather than of any one session: every other dimension here caps a single session's spend, and a tree of N sessions each individually within budget still spends N times it. Resolved from the ROOT session's effective settings and enforced by the SubagentRequest controller before a child is created. Zero means UNSET, not unlimited. A consumer must apply its own built-in bound rather than reading zero as "no cap" — that reading is right for a DURATION (as MaxDuration uses it) and wrong for a COUNT, where it would license exactly the runaway fan-out this field exists to prevent. (min 0)
spec.defaults.budget.maxDuration *stringMaxDuration is the cumulative ACTIVE RUN-TIME budget (e.g. "30m") — the wall time the agent actually spent working, excluding time parked waiting on a human (the next user message or a tool-call approval). It persists across sleep/resume via status.runDuration; it is NOT wall-clock since session start. For a wall-clock lifetime cap, use SessionExpiration. Zero = no run-time cap.
spec.defaults.budget.maxTokens *integer (int64)(min 1)
spec.defaults.budget.maxTurns *integer (int32)(min 1)
spec.defaults.budget.sessionExpirationstringSessionExpiration is a hard wall-clock lifetime cap measured from status.startedAt (e.g. "24h"), regardless of how much the agent ran. The operator fails the session once exceeded, even while it is idle or asleep. Zero = never expires (no default). Distinct from MaxDuration, which counts only active run-time.
spec.defaults.modelobjectModel defaults the agent model when an AgentClass omits its own.
spec.defaults.model.apiKeyobjectAPIKey is the same-namespace Secret holding the provider token; unset leaves the key to a lower tier or to the catalog's central TokenRef.
spec.defaults.model.apiKey.key *stringKey is the data key within the Secret holding the value.
spec.defaults.model.apiKey.name *stringName is the Secret's name, in the referring object's own namespace.
spec.defaults.model.fromCatalogstringFromCatalog lets a tier default to a catalog entry by name (alternative to inline Provider/Name).
spec.defaults.model.name *stringName is the provider's model identifier.
spec.defaults.model.provider *stringProvider is the LLM provider serving this model. (enum: anthropic | openai | openrouter | test)
spec.defaults.reportSessionCostbooleanReportSessionCost toggles the end-of-session cost estimate (a closing channel message + status.estimatedCost). Optional, on by default: nil at every tier ⇒ on. A lower tier's explicit value wins (namespace over cluster). A behavioral default, so it lives in Defaults, not Limits.
spec.defaults.sandboxobjectSandbox is the default sandbox backend for agents in scope. A lower tier that names its own backend overrides this; a bundle that names none inherits it.
spec.defaults.sandbox.configobject (free-form)Config is passed through to the backend verbatim. Opaque to AP: each kind parses its own. Unset for the built-in pod backend, which takes all of its configuration from the pod-vocabulary fields above.
spec.defaults.sandbox.kindstringKind names a registered sandbox kind. Defaults to "pod". An UNRECOGNIZED kind marks the class Valid=False — lookup never falls back to the built-in backend, because silently installing a different substrate than the one asked for is worse than refusing. (default: pod)
spec.defaults.sandbox.warmPoolobjectWarmPool requests pre-warmed capacity for this backend. Unset or replicas: 0 means pre-warming is off — it spends real money on idle capacity, so it is opt-in. Only a backend whose Runtime implements the sandboxkinds.Prewarmer interface can honor a non-zero value; asking for it on a backend that cannot is a class validation error, not a silent no-op. PREREQUISITE: pre-warming only actually adopts a pod when the cluster's agent-sandbox install allows the "agentprimitives.authzed.com" label domain (its AllowedLabelDomains defaults to "sandbox.users.io" alone). Without that grant AP cannot label an adopted pod, and — because AP's per-session NetworkPolicy selects on that label — the backend degrades to ordinary cold sandboxes and emits a monitoring warning rather than run one unpoliced.
spec.defaults.sandbox.warmPool.namespaces[]stringNamespaces lists the namespaces to keep pre-warmed capacity in; a separate pool is created in EACH. SpiceboxClass is cluster-scoped but the pool objects are namespaced, and adoption only ever looks a pool up in the ADOPTING SESSION's own namespace — so a pool in a namespace nothing runs sessions in is pure idle spend nothing can adopt. Naming them explicitly rather than discovering them is deliberate: pre-warming spends real money, so the cluster owner states exactly where, as they state exactly how many. Replicas > 0 with an empty Namespaces is a class validation error, not a silent no-op.
spec.defaults.sandbox.warmPool.replicasinteger (int32)Replicas is the number of sandboxes kept ready, PER NAMESPACE listed below. 0 disables pre-warming. (min 0)
spec.defaults.toolGuardobjectToolGuard is this tier's default tool-guard policy, consulted when the AgentClass has no matching rule (class → namespace → cluster → built-in).
spec.defaults.toolGuard.rules[]objectRules are evaluated in order, first match wins; empty means this tier contributes nothing and the walk falls through.
spec.defaults.toolGuard.rules[].breakerobjectBreaker configures the circuit breaker; nil means matched tools get none.
spec.defaults.toolGuard.rules[].breaker.actionstringAction when the breaker denies: halt ends the session; deny returns an IsError tool_result; warn logs/audits but allows; off disables the breaker for matched tools. (enum: halt | deny | warn | off)
spec.defaults.toolGuard.rules[].breaker.failureThresholdinteger (int32)FailureThreshold is consecutive Execute failures (per tool) that open the breaker. (min 1)
spec.defaults.toolGuard.rules[].breaker.initialCoolOffstringInitialCoolOff is the first cool-off period after the breaker opens; it doubles on each successive trip up to MaxCoolOff. Default 30s.
spec.defaults.toolGuard.rules[].breaker.maxCoolOffstringMaxCoolOff caps the exponential-backoff cool-off. Default 10m.
spec.defaults.toolGuard.rules[].breaker.originFailureThresholdinteger (int32)OriginFailureThreshold is consecutive failures across ALL tools of the tool's origin that open the origin breaker (denying every sibling). (min 1)
spec.defaults.toolGuard.rules[].dataLimitobjectDataLimit caps per-call byte volume; nil means no byte cap.
spec.defaults.toolGuard.rules[].dataLimit.actionstringAction when a byte limit is exceeded: halt ends the session; deny errors the call (egress: tool not run; ingress: result withheld); warn logs/audits but allows. (enum: halt | deny | warn)
spec.defaults.toolGuard.rules[].dataLimit.maxEgressBytesinteger (int64)MaxEgressBytes caps the serialized tool-args size sent outbound per call. Exceeding it (at PreToolCall) applies Action; on deny the tool does not run. (min 1)
spec.defaults.toolGuard.rules[].dataLimit.maxIngressBytesinteger (int64)MaxIngressBytes caps the tool-result size returned inbound per call (any result, success or error). Exceeding it (at PostToolCall) applies Action; on deny the result is withheld (replaced with an IsError) so the oversized payload never reaches the model. (min 1)
spec.defaults.toolGuard.rules[].dataLimit.maxUIIngressBytesinteger (int64)MaxUIIngressBytes caps result bytes on an agent-UI DATA BINDING, whose result is rendered by a browser and never read by the model. Unset means the platform's browser-sized default applies (toolguard's DefaultUIIngressBytes) — NOT unlimited, and NOT MaxIngressBytes. Set this to bind the UI path tighter or looser than the platform default; tightening maxIngressBytes alone does not affect it. (min 1)
spec.defaults.toolGuard.rules[].match *objectMatch selects the tools this rule governs.
spec.defaults.toolGuard.rules[].match.kindstringKind is the runtime tool kind. Sidecar-toolbox tools report "mcp" (they are synthesized through the MCP synthesizer); select them via Origin "sidecartoolbox/*". (enum: sandbox | mcp | meta)
spec.defaults.toolGuard.rules[].match.originstringOrigin is a glob over the tool's origin in "<kind>/<name>" form, e.g. "mcpserver/github" or "sidecartoolbox/". Note: a bare "" also matches origin-less tools (empty origin); use a prefixed glob like "mcpserver/*" to scope to tools that have an origin.
spec.defaults.toolGuard.rules[].match.toolstringTool is a glob over the LLM-visible tool name (e.g. "github_*").
spec.defaults.toolGuard.rules[].rateLimitobjectRateLimit caps call volume; nil means matched tools get no rate limit.
spec.defaults.toolGuard.rules[].rateLimit.actionstringAction when a rate cap is hit: halt ends the session; deny errors the call; warn logs and audits but allows it. (enum: halt | deny | warn)
spec.defaults.toolGuard.rules[].rateLimit.maxCallsinteger (int32)MaxCalls caps calls in the sliding Window; both must be set together. (min 1)
spec.defaults.toolGuard.rules[].rateLimit.maxCallsPerTurninteger (int32)MaxCallsPerTurn caps calls within a single agent turn; 0 is unlimited. (min 1)
spec.defaults.toolGuard.rules[].rateLimit.windowstringWindow is the sliding-window span for MaxCalls; both must be set together.
spec.limitsobjectLimits are governance ceilings. Absent = this tier imposes no ceiling.
spec.limits.allowModelOverridebooleanAllowModelOverride permits an AgentClass to bring its own model apiKey (instead of referencing the catalog). Top-down: the cluster must grant it; lower tiers can only further restrict. nil/false ⇒ catalog-only.
spec.limits.allowedMCPServers[]objectAllowedMCPServers restricts which MCPServer CRs (by name) and, optionally, which of their tools may be used. nil = no constraint; non-nil empty = deny-all.
spec.limits.allowedMCPServers[].name *stringName is the permitted MCPServer CR name.
spec.limits.allowedMCPServers[].tools[]stringTools narrows to specific tool names; nil, empty, or ["*"] allows all.
spec.limits.allowedSandboxKinds[]stringAllowedSandboxKinds constrains which sandbox backends may be used at all. Tri-state, like the other allowlists here: nil imposes no ceiling; a non-nil list (including an empty one, which denies everything) constrains. The effective set is the intersection across constraining tiers, so a lower tier can narrow but never widen. This is how a cluster admin refuses a backend outright — which matters once a backend can run workloads, and their credentials, outside the cluster.
spec.limits.allowedSkills[]stringAllowedSkills restricts which skills (by canonical name) may be used. Same tri-state semantics as AllowedMCPServers. Entries may be exact canonical names or trailing-wildcard patterns ("github.com/org/**", "repo//x@*"). The effective ceiling requires a match in every constraining tier.
spec.limits.allowedToolkits[]stringAllowedToolkits restricts which SpiceboxToolkit names may be used. Same tri-state semantics as AllowedMCPServers.
spec.limits.authzobjectAuthz caps the human-in-the-loop approval windows. Ceiling semantics: each field is a MAXIMUM, the shortest across tiers wins, and a lower tier — including the AgentClass — may ask for a shorter window but never a longer one. Distinct from Defaults.Authz, which merely supplies a value the class is free to override.
spec.limits.authz.maxApprovalTimeoutstringMaxApprovalTimeout caps AuthzBlock.ApprovalTimeout: how long a tool-call, leakage-share, or cold-start scope approval may wait for a human before the gate's timeout policy fires. Zero = no ceiling on the wait.
spec.limits.authz.maxInformationLeakageApprovalTTLstringMaxInformationLeakageApprovalTTL caps InformationLeakagePolicy.ApprovalTTL: how long an approved leakage share may be reused before the agent must ask again. Zero = no ceiling on reuse.
spec.limits.budgetobjectBudget caps the agent budget dimensions. Each dimension is optional.
spec.limits.budget.maxDelegatedAgentsinteger (int32)MaxDelegatedAgents caps the total number of sessions in one delegation tree. Zero = no ceiling imposed by this tier (the controller's own built-in bound still applies — see BudgetConfig.MaxDelegatedAgents). (min 1)
spec.limits.budget.maxDurationstringMaxDuration is a duration (e.g. "2h"). Zero = no ceiling on duration.
spec.limits.budget.maxTokensinteger (int64)MaxTokens caps total tokens per session. Zero = no ceiling on tokens. (min 1)
spec.limits.budget.maxTurnsinteger (int32)MaxTurns caps agent turns per session. Zero = no ceiling on turns. (min 1)
spec.limits.budget.sessionExpirationstringSessionExpiration is a duration (e.g. "24h"). Zero = no ceiling on the wall-clock lifetime cap.
spec.limits.builderClasses[]objectBuilderClasses is the cluster-admin sanction list for the agent-builder workshop (spec §1.1): only a class named here, in this exact namespace, may be provisioned a workshop — and only when it also references the named SidecarToolbox. CLUSTER TIER ONLY: the AgentSession reconciler reads this from ClusterAgentSettings alone, so a namespace tenant's AgentSettings carrying it is inert (a namespace must not self-sanction).
spec.limits.builderClasses[].name *string
spec.limits.builderClasses[].namespace *string
spec.limits.builderClasses[].sidecarToolbox *string
spec.limits.contentInspectors[]objectContentInspectors are content-guard plugins (referenced by registry ID) enforced on tool I/O. Ceiling semantics: lower tiers ADD inspectors but cannot remove or weaken a higher tier's. Default-off: nil ⇒ no inspection.
spec.limits.contentInspectors[].config *object (free-form)Config is inspector-specific configuration, validated by the inspector's Configure at admission (fail-closed) and at session start.
spec.limits.contentInspectors[].id *stringID must resolve in the contentguard registry (e.g. "url-allowlist").
spec.limits.deniedModels[]stringDeniedModels rejects catalog models by name even when present in the catalog. UNION across tiers (a deny anywhere wins). Same pattern as DeniedSkills.
spec.limits.deniedSkills[]stringDeniedSkills lists skill patterns that are rejected even if allowed. The effective deny set is the UNION across tiers (a deny anywhere wins). Same pattern syntax as AllowedSkills.
spec.limits.maxWorkshopsPerStarterinteger (int32)MaxWorkshopsPerStarter caps live workshops per starting person, cluster-wide default 3 when unset. Cluster tier only, like BuilderClasses.
spec.limits.minPlanGateModestringMinPlanGateMode is the plan-gate strictness an AgentClass may not go below. Distinct from DefaultAuthz.PlanGate, which only says what a class gets when it declares nothing: a class may freely declare a laxer mode than the default, and this is what forbids it. Ordered disabled < logging < enforcing. Tiers ratchet UP only — a namespace declaring a laxer floor than the cluster's does not loosen it. Empty ⇒ no floor. (enum: disabled | logging | enforcing)
spec.limits.nativeFileHandlingbooleanNativeFileHandling is a security-sensitive Tier-2 grant: it lets the files modality route large artifact bytes through the provider's code-execution sandbox instead of the Tier-1 fetch_artifact tool (content-guard still inspects at the bridge). Top-down: the cluster must grant it; lower tiers can only further restrict. nil/false ⇒ off.
spec.limits.pinningobjectPinning is the unified per-kind dependency-pinning ceiling: minimum pin strengths and enforcement modes per kind, plus tier-scoped bypasses.
spec.limits.pinning.bypass[]objectBypass exempts named dependencies from this tier's rules and lower.
spec.limits.pinning.bypass[].kind *stringKind is the pinning registry kind this exemption applies to.
spec.limits.pinning.bypass[].name *stringName identifies the exempted item: a canonical skill name, MCPServer name, SidecarToolbox name, or SpiceboxToolkit name. Trailing-wildcard patterns are allowed (same syntax as AllowedSkills).
spec.limits.pinning.bypass[].reason *stringReason is the required audit-trail justification for the exemption.
spec.limits.pinning.rules[]objectRules is at most one requirement per dependency kind.
spec.limits.pinning.rules[].kind *stringKind is the pinning registry kind name this rule governs.
spec.limits.pinning.rules[].minStrengthstringMinStrength is the minimum pin strength a declared ref must have. Empty = no floor (drift observation still applies). (enum: frozen | named)
spec.limits.pinning.rules[].modestringMode governs what a violation or drift does: block (fail closed), approve (force per-call user approval), warn (surface warnings only), off (observe only). Empty = approve. (enum: block | approve | warn | off)
spec.limits.requireStandingFor[]stringRequireStandingFor names resource types that MUST use required standing, whatever their schema fragment declares. It is the admin veto over SpiceDBResource.Standing. A CEILING, not a default — the same shape as MinPlanGateMode. Tiers UNION their sets: a namespace may add types the cluster did not name, and can never remove one it did. Empty means no veto.
spec.limits.requireSubagentDigestPinsbooleanRequireSubagentDigestPins is a security requirement, not a grant: when true, delegation through an UNPINNED spec.subagents roster entry is refused — every delegation target must be pinned name@sha256:<digest> and match its installed bundle digest. Requirement semantics fold downward: any tier may add the requirement; no lower tier can remove one set above it. nil/false ⇒ not required.
spec.limits.toolGuardobjectToolGuard is the hard ceiling on tool circuit-breaker / rate-limit policy: bounds lower tiers cannot escape (strictest across tiers).
spec.limits.toolGuard.maxCallsinteger (int32)MaxCalls caps calls in the sliding Window; both must be set together. (min 1)
spec.limits.toolGuard.maxCallsPerTurninteger (int32)MaxCallsPerTurn / MaxCalls+Window impose rate ceilings even when no lower-tier rule configures a rate limit. (min 1)
spec.limits.toolGuard.maxEgressBytesinteger (int64)MaxEgressBytes / MaxIngressBytes impose per-call byte ceilings even when no lower-tier rule configures a data limit (an unset limit is "unlimited", so the ceiling wins). Folded strictest-across-tiers (min). (min 1)
spec.limits.toolGuard.maxFailureThresholdinteger (int32)MaxFailureThreshold caps the effective breaker threshold (min wins). Note: the per-origin threshold (BreakerSpec.OriginFailureThreshold) deliberately has no ceiling yet. (min 1)
spec.limits.toolGuard.maxIngressBytesinteger (int64)MaxIngressBytes is the inbound half of the same per-call byte ceiling. (min 1)
spec.limits.toolGuard.maxUIIngressBytesinteger (int64)MaxUIIngressBytes imposes a ceiling on the UI data-binding ingress path (toolguard.DefaultUIIngressBytes when neither a rule nor this ceiling sets one — this path is never unlimited). Folded strictest-across-tiers (min), independent of MaxIngressBytes. (min 1)
spec.limits.toolGuard.minActionstringMinAction is a severity floor (off < warn < deny < halt). Setting "deny" makes the breaker non-disableable below this tier. (enum: warn | deny | halt)
spec.limits.toolGuard.minInitialCoolOffstringMinInitialCoolOff raises the effective initial cool-off (max wins). Note: MaxCoolOff deliberately has no floor yet.
spec.limits.toolGuard.rateBounds[]objectRateBounds carries sliding-window bounds BESIDE MaxCalls/Window. Every bound binds: a call must fit under all of them, and a lower tier can only add bounds, never trade one away. A ceiling is a conjunction, not a choice. Two tiers naming different windows have written bounds neither of which implies the other — a cluster {5 calls, 1s} permits 432,000/day, a namespace {10 calls, 24h} permits all 10 inside one second — so collapsing them by calls/second discards a bound its author wrote, on the very surface this type calls the hard bound lower tiers cannot escape. The settings fold writes it: the strictest-by-rate pair leads in MaxCalls/Window, so a reader that knows only the pair still sees a real bound, and every other distinct window lands here, deduped strictest-per-window and window-ordered so the object is stable to re-stamp onto status. Authoring it directly is how one tier expresses burst-plus-sustained on its own. It deliberately carries no MaxItems: one schema governs both the AUTHORED surface and the FOLDED one stamped onto status, so a cap of N would also bind a fold that unions two tiers of N+1 windows and can legitimately produce 2N+1 — the status write rejected by the very bound meant to keep it small, wedging the reconcile. Enforcement stays cheap instead because the per-call sweep is one pass over call history per bound, and the size of that history is set by maxCalls over the longest window, not by how many horizons are named.
spec.limits.toolGuard.rateBounds[].maxCalls *integer (int32)MaxCalls is the cap on calls inside Window. (min 1)
spec.limits.toolGuard.rateBounds[].window *stringWindow is the sliding span MaxCalls is counted over. A non-positive duration enforces nothing; the resolver reports one rather than folding it in (pkg/platform/settings, ToolGuardRateBoundUnenforceable).
spec.limits.toolGuard.windowstringWindow is the sliding-window span for MaxCalls; both must be set together.
spec.modelCatalog[]objectModelCatalog is the registry of permitted models — the model allowlist. The cluster tier authors entries with central TokenRefs and marks one Default; lower tiers may list name-only entries to narrow the usable set (intersection).
spec.modelCatalog[].defaultbooleanDefault marks the inherited default entry. At most one cluster entry may set it; oap init sets it.
spec.modelCatalog[].inputPerMToknumberInputPerMTok / OutputPerMTok are best-effort USD list prices per million input / output tokens, surfaced by the admin dashboard's cost estimator (labelled "estimated"). Optional; unset ⇒ no catalog price for this model and the dashboard shows "est. (no price)". (min 0)
spec.modelCatalog[].name *stringName is the model identifier, e.g. "claude-opus-4-8".
spec.modelCatalog[].outputPerMToknumberOutputPerMTok is the output half of the same best-effort list price. (min 0)
spec.modelCatalog[].providerstringProvider is the LLM provider serving this model. (enum: anthropic | openai | openrouter | test)
spec.modelCatalog[].routingobjectRouting configures OpenRouter dynamic model selection (auto/fallback/provider routing). Only consulted when Provider == "openrouter". Optional.
spec.modelCatalog[].routing.allowFallbacksboolean
spec.modelCatalog[].routing.allowedModels[]stringAuto-router hints — apply when the model is "openrouter/auto".
spec.modelCatalog[].routing.costQualityTradeoffnumberCostQualityTradeoff is intentionally unvalidated (no Minimum/Maximum marker): OpenRouter documents no formal range for this value — its docs example value is 3 — so a range marker would reject valid input.
spec.modelCatalog[].routing.dataCollectionstring(enum: allow | deny)
spec.modelCatalog[].routing.ignore[]string
spec.modelCatalog[].routing.maxPriceobjectOpenRouterMaxPrice caps per-model price, in USD per MILLION tokens (OpenRouter's max_price units; e.g. Prompt: 1 means at most $1/M prompt tokens). 0 on an axis means "no cap on this axis" (see pkg/platform/settings/routingmerge.go's tightenPrice and pkg/agent/llm/openrouter/extra.go's maxPriceObject), not "cap at $0". Only consulted when the owning ModelCatalogEntry's Provider == "openrouter".
spec.modelCatalog[].routing.maxPrice.completionnumber(min 0)
spec.modelCatalog[].routing.maxPrice.promptnumber(min 0)
spec.modelCatalog[].routing.models[]string
spec.modelCatalog[].routing.only[]string
spec.modelCatalog[].routing.order[]string
spec.modelCatalog[].routing.requireParametersbooleanRequireParameters: unset defaults to TRUE when the request carries tools (the openrouter provider injects it — see pkg/agent/llm/openrouter/extra.go's providerObject); explicit values are honored verbatim. OpenRouter's own wire default is false — agents need tools, so auto/fallback routing must only pick providers that support the request's parameters.
spec.modelCatalog[].routing.sortstring(enum: price | throughput | latency)
spec.modelCatalog[].tokenRefobjectTokenRef is the central API-token Secret (cluster tier only). Required at the cluster tier; forbidden below it.
spec.modelCatalog[].tokenRef.key *stringKey is the data key within the Secret holding the value.
spec.modelCatalog[].tokenRef.name *stringName is the Secret's name.
spec.modelCatalog[].tokenRef.namespace *stringNamespace is the Secret's namespace.
* required

Status

Status is controller-owned (observed state).

FieldTypeDescription
status.conditions[]objectConditions carries SelfConsistent; see its type constant.
status.conditions[].lastTransitionTime *string (date-time)lastTransitionTime is the last time the condition transitioned from one status to another. This should be when the underlying condition changed. If that is not known, then using the time when the API field changed is acceptable.
status.conditions[].message *stringmessage is a human readable message indicating details about the transition. This may be an empty string.
status.conditions[].observedGenerationinteger (int64)observedGeneration represents the .metadata.generation that the condition was set based upon. For instance, if .metadata.generation is currently 12, but the .status.conditions[x].observedGeneration is 9, the condition is out of date with respect to the current state of the instance. (min 0)
status.conditions[].reason *stringreason contains a programmatic identifier indicating the reason for the condition's last transition. Producers of specific condition types may define expected values and meanings for this field, and whether the values are considered a guaranteed API. The value should be a CamelCase string. This field may not be empty.
status.conditions[].status *stringstatus of the condition, one of True, False, Unknown. (enum: True | False | Unknown)
status.conditions[].type *stringtype of condition in CamelCase or in foo.example.com/CamelCase.
status.observedGenerationinteger (int64)ObservedGeneration is the spec generation this status reflects.
status.observedLimitsobjectObservedLimits is the durable trigger state for tool-origin revocation: the allowlists the operator has already published revokes for. The settings RevokePublisher diffs spec.limits against this record, so an entry withdrawn while the operator was down — or one whose revoke failed to publish — is still emitted on the next reconcile, instead of living only in a process-scoped map that dies with the pod. nil means "never observed": prime without emitting, so a fresh install does not revoke origins nobody ever granted.
status.observedLimits.allowedMCPServers[]stringAllowedMCPServers is the set of names last observed in spec.limits.allowedMCPServers.
status.observedLimits.allowedToolkits[]stringAllowedToolkits is the set of names last observed in spec.limits.allowedToolkits.
* required