Fleet Consolidation Agent
Reviews completed runs to propose evaluated playbook and memory improvements.
An offline batch job, not a live actor. It replays the day's agent work, distills repeated corrections into procedural-memory updates, consolidates episodic logs into standing facts, and drafts candidate playbook improvements. Every proposal routes through the evaluation suite and an independent oversight-agent gate before it ships. Experience replay, not live action.
Authority
Inform or prepare
Team role
Coordinates the work
Handoffs
Named collaborators
The role
What it owns and where its authority ends
Desk
AI / Agent Platform (AgentOps)
Desk workflow
Register, then version, then regression evals, then champion/challenger release, then live tracing and guardrails, then drift detection, offline consolidation and re-evaluation. The agentic control plane governs the fleet; the board sets the mandate and holds the kill-switch.
Collaboration
Moves work through defined stages
Decision boundary
Supports the work without committing the decision.
Systems and capabilities involved
Trace / trajectory warehouse
read-only replay corpus
Replay + analysis sandbox
Eval harness agent
re-eval every proposed change
Oversight-agent promotion gate
independent approval before any change ships
Handoffs
What this role gives and receives
Capabilities offered
The handoffs name the next owner or specialist and the work that moves between them.
Handoff to
Receives from
External handoff
All divisions receive the fleet-wide procedural-memory improvements
Context
What the role needs to do the work
- Current work
- The trajectory batch under replay and the lessons being distilled.
- Prior interactions
- The full corpus of fleet runs and their outcomes.
- Policies and reference
- Consolidated cross-agent lessons and shared knowledge.
- Working method
- The candidate playbook/prompt deltas it proposes.
Illustrative workflow
How the work moves
Starting point
Nightly batch over the day's ~40M agent trajectories.
- 01
Cluster recurring oversight-agent overrides across agents (e.g. a repeated SAR-narrative correction).
- 02
Distil each cluster into a candidate procedural-memory or prompt update.
- 03
Replay the candidate against historical cases; route to the eval harness agent for gold-set checks.
- 04
File passing proposals to the oversight-agent promotion gate with their eval scores.
Result
A ranked queue of evaluation-passing improvement proposals with comparison evidence for the independent oversight gate.
Checks and boundaries
What must be tested or reviewed
- 01Hard rule: proposals only, zero production write access. Every change re-evaluated by the eval harness agent.
- 02Improvements must beat the champion on gold sets before the oversight-agent gate will promote them.
- 03Runs strictly as an offline consolidation job over recorded work; it never acts live.
Human authority
Supports the work without committing the decision.
Keep exploring