Eval Harness Agent
Runs regression, independent-review and adversarial checks before an agent release.
Gates every agent release. It runs gold-set regression suites, has independent judge agents score outputs against rubrics, fires adversarial red-team cases, and runs champion/challenger comparisons before any new version reaches production. A drop on any safety-critical suite blocks the release.
Authority
Act within policy
Team role
Provides independent challenge
Handoffs
Named collaborators
The role
What it owns and where its authority ends
Desk
AI / Agent Platform (AgentOps)
Desk workflow
Register, then version, then regression evals, then champion/challenger release, then live tracing and guardrails, then drift detection, offline consolidation and re-evaluation. The agentic control plane governs the fleet; the board sets the mandate and holds the kill-switch.
Collaboration
Coordinates specialist contributions
Decision boundary
Acts only within an explicit policy, permission and escalation boundary.
Systems and capabilities involved
Release pipeline
gate releases on eval pass
Gold-set + rubric store
Independent judge agents
Scoring + statistics sandbox
Red-team prompt generator
Handoffs
What this role gives and receives
Capabilities offered
The handoffs name the next owner or specialist and the work that moves between them.
Handoff to
Receives from
Receives from
Context
What the role needs to do the work
- Current work
- The candidate version under test and its running scorecard.
- Prior interactions
- Every prior eval run per agent: the full regression history.
- Policies and reference
- Gold sets, rubrics, red-team attack library, pass thresholds.
- Working method
- Which eval suites matter for which agent class.
Illustrative workflow
How the work moves
Starting point
A new version of the Sanctions Disposition Agent is submitted for release.
- 01
Run the gold-set regression suite; spawn judge workers to score dispositions.
- 02
Fire the red-team battery (alias-evasion, prompt-injection in counterparty names).
- 03
Run champion/challenger against the live version on held-out true-match cases.
- 04
Compile the scorecard and return it to the submitting team.
Result
Release blocked: the challenger improved false-positive release but missed one gold true-match, a hard fail on a safety-critical suite. Scorecard and the failing case attached for the author.
Checks and boundaries
What must be tested or reviewed
- 01The eval rig evals itself: judge calibration checked against the board-anchored ground-truth gold set.
- 02Champion/challenger required before any version promotion; no silent swaps.
- 03Regression gate is hard: a drop on any safety-critical suite blocks the release.
- 04Red-team suite refreshed continuously from new jailbreak/abuse patterns.
Human authority
Acts only within an explicit policy, permission and escalation boundary.
Keep exploring