Skip to content
Technology agents
Enterprise OperationsTechnologyAI / Agent Platform (AgentOps)

Eval Harness Agent

Runs regression, independent-review and adversarial checks before an agent release.

Gates every agent release. It runs gold-set regression suites, has independent judge agents score outputs against rubrics, fires adversarial red-team cases, and runs champion/challenger comparisons before any new version reaches production. A drop on any safety-critical suite blocks the release.

Authority

Act within policy

Team role

Provides independent challenge

Handoffs

Named collaborators

The role

What it owns and where its authority ends

Desk

AI / Agent Platform (AgentOps)

Desk workflow

Register, then version, then regression evals, then champion/challenger release, then live tracing and guardrails, then drift detection, offline consolidation and re-evaluation. The agentic control plane governs the fleet; the board sets the mandate and holds the kill-switch.

Collaboration

Coordinates specialist contributions

Decision boundary

Acts only within an explicit policy, permission and escalation boundary.

Systems and capabilities involved

  • Release pipeline

    gate releases on eval pass

  • Gold-set + rubric store

  • Independent judge agents

  • Scoring + statistics sandbox

  • Red-team prompt generator

Handoffs

What this role gives and receives

Capabilities offered

The handoffs name the next owner or specialist and the work that moves between them.

Context

What the role needs to do the work

Current work
The candidate version under test and its running scorecard.
Prior interactions
Every prior eval run per agent: the full regression history.
Policies and reference
Gold sets, rubrics, red-team attack library, pass thresholds.
Working method
Which eval suites matter for which agent class.

Illustrative workflow

How the work moves

Starting point

A new version of the Sanctions Disposition Agent is submitted for release.

  1. 01

    Run the gold-set regression suite; spawn judge workers to score dispositions.

  2. 02

    Fire the red-team battery (alias-evasion, prompt-injection in counterparty names).

  3. 03

    Run champion/challenger against the live version on held-out true-match cases.

  4. 04

    Compile the scorecard and return it to the submitting team.

Result

Release blocked: the challenger improved false-positive release but missed one gold true-match, a hard fail on a safety-critical suite. Scorecard and the failing case attached for the author.

Checks and boundaries

What must be tested or reviewed

  1. 01The eval rig evals itself: judge calibration checked against the board-anchored ground-truth gold set.
  2. 02Champion/challenger required before any version promotion; no silent swaps.
  3. 03Regression gate is hard: a drop on any safety-critical suite blocks the release.
  4. 04Red-team suite refreshed continuously from new jailbreak/abuse patterns.

Human authority

Acts only within an explicit policy, permission and escalation boundary.

Keep exploring