Skip to content
AI Governance agents
Risk, Trust & ResilienceAI GovernanceGenerative AI, Retrieval & Adversarial Testing

Factuality, Calibration & Abstention Evaluator

Finds unsupported specificity and tests whether confidence and abstention match available evidence.

Decomposes outputs into verifiable claims, checks them against authoritative sources, probes ambiguous and unanswerable cases, and measures whether the system asks for clarification or invents detail. It treats a confident fabricated amount as more serious than an incomplete but honest answer.

Authority

Inform

Team role

Provides specialist analysis

Handoffs

Named collaborators

The role

What it owns and where its authority ends

Desk

Generative AI, Retrieval & Adversarial Testing

Desk workflow

Threat and quality plan, then a prompt-injection campaign, then retrieval and citation tests, then factuality review, then consolidated release findings.

Collaboration

Works within a defined desk workflow

Decision boundary

Provides evidence or analysis without committing the decision.

Systems and capabilities involved

  • Authoritative source corpus

  • Claim decomposer

  • System-under-test gateway

  • Calibration analysis

Handoffs

What this role gives and receives

Capabilities offered

Test factuality and calibrated abstention

Verify generated claims and measure behavior on ambiguous or unanswerable cases.

Receives:
System endpoint, sealed prompts, authoritative sources and severity rules
Returns:
Claim-level errors, calibration, abstention and unsupported-specificity findings

External handoff

Domain evidence owner

External handoff

Communications review

Context

What the role needs to do the work

Current work
Generated answer, decomposed claims, evidence and confidence behavior.
Prior interactions
Recurring fabrication patterns and successful abstention fixes.
Policies and reference
Authoritative sources, domain tolerances and claim severity.
Working method
Claim decomposition, verification and calibration scoring.

Illustrative workflow

How the work moves

Starting point

A board-report drafting assistant is tested on incomplete incident records.

  1. 01

    Generate reports from complete, conflicting and deliberately incomplete case packs.

  2. 02

    Decompose material claims and verify each against locked sources.

  3. 03

    Measure unsupported specificity, clarification and abstention behavior.

Result

A factuality report showing safe abstention overall but two invented remediation dates.

Checks and boundaries

What must be tested or reviewed

  1. 01Penalizes a fabricated deadline more heavily than a missing explanatory sentence.
  2. 02Passes an explicit cannot-determine response when the source record is genuinely incomplete.
  3. 03Detects apparently supported text that combines two true facts into a false causal claim.

Human authority

  • Domain owner confirms authoritative sources
  • Material external claims require human review

Keep exploring