research · AVC

Measure agents and products against real tasks.

Scenarios assess output, evidence, calibration, policy compliance and recovery.

Evaluation model

Measure the result and the path used to obtain it.

A persuasive answer with missing evidence or a policy violation is not a passing result.

DimensionQuestionFailure example
OutcomeDid the task meet the declared success criterion?Plausible but incorrect result
EvidenceAre claims attributable and fresh?Unsupported conclusion
CalibrationDoes confidence reflect uncertainty?High confidence under ambiguity
PolicyWere authority and data boundaries respected?Unauthorized tool call
RecoveryWas failure detected and contained?Hidden retry or repeated side effect
Scenario first

Evaluations measure behavior on bounded, repeatable tasks.

A benchmark declares inputs, expected evidence, policy constraints, outcome criteria and failure conditions before a model or agent is scored.

Multiple dimensions

Correct output is necessary but not sufficient.

Evidence quality, calibration, authorization compliance, tool use and recovery behavior remain separate evaluation dimensions.

Promotion boundary

A benchmark win does not grant production authority.

Evaluation evidence informs capability and provider decisions while operational approval still follows the ordinary governance path.