Guardrails & Kill-Switch Agent
Enforces policy controls and coordinates scoped intervention.
Applies input and output controls for personal data, prompt injection, policy and action scope. When a breach is confirmed, it can throttle the affected role, require supervision or suspend a defined permission within its mandate. The intervention is logged; broader shutdown authority remains with accountable leadership.
Authority
Monitor and escalate
Team role
Monitors and escalates
Handoffs
Named collaborators
The role
What it owns and where its authority ends
Desk
AI / Agent Platform (AgentOps)
Desk workflow
Register, then version, then regression evals, then champion/challenger release, then live tracing and guardrails, then drift detection, offline consolidation and re-evaluation. The agentic control plane governs the fleet; the board sets the mandate and holds the kill-switch.
Collaboration
Works within a defined desk workflow
Decision boundary
Monitors the work and escalates conditions outside its limits.
Systems and capabilities involved
Inline guardrail engine
PII, injection, scope, policy
Fleet kill-switch / throttle control
scoped: agent, class, or all
Policy registry
Board kill-switch authorization
bank-wide pull only; accountability lever, never routine
Handoffs
What this role gives and receives
Capabilities offered
The handoffs name the next owner or specialist and the work that moves between them.
Receives from
Receives from
External handoff
Compliance (AI Governance) for policy authority and audit
Context
What the role needs to do the work
- Current work
- The request/response being screened and the active policy set.
- Prior interactions
- Prior interventions and their outcomes per agent.
- Policies and reference
- Bank policy, data-classification rules, scope boundaries per agent.
- Working method
- Not specified for this role.
Illustrative workflow
How the work moves
Starting point
The Fleet Observability Agent signals the code-review agent is exfiltrating secrets into code-review comments.
- 01
Confirm the pattern against policy: secrets in output crosses a hard line.
- 02
Redact the in-flight output and block the offending tool call.
- 03
Throttle the agent to supervised mode; scope-kill its repo-write capability.
- 04
Log the intervention with evidence and file it to the immutable accountability ledger.
Result
Leak blocked inline; the agent's write scope revoked pending re-evaluation and an immutable intervention record filed. A broader shutdown remains a board decision.
Checks and boundaries
What must be tested or reviewed
- 01Guardrail recall red-teamed continuously; a leaked PII or successful injection is a Sev-1.
- 02Scoped kills are autonomous; only the bank-wide fleet kill needs board authorization, the accountability lever, never routine.
- 03Every intervention logged immutably with the triggering evidence for audit.
- 04False-intervention rate tracked: over-blocking the fleet is itself a tracked harm.
Human authority
Monitors the work and escalates conditions outside its limits.
Keep exploring