Evaluation and assurance
PUBLISHEDEstablishing with evidence whether a system is good enough, and when it stops being.
| Excluded | Where it lives |
|---|---|
| Formal audit engagement, independence and reporting | O3 |
| Research-depth statistical methodology | Out of framework |
| Setting the criteria themselves | D5 |
Where this domain abuts another, and how the line is drawn
D5 defines what good looks like. D6 establishes whether it was achieved.
D6 produces assurance evidence. O3 covers the professional obligations of the person auditing it.
D6 detects a security failure as a quality failure. D8 covers the adversary that caused it.
Every identifier is a permanent address. Indicators are normative; they state what would be observed in a person who meets the statement.
L1 Aware
4 statementsDescribe why AI output requires evaluation before it is relied upon.
Distinguish a system that has been demonstrated from one that has been evaluated.
State what a test case is and why evaluation requires more than a small number of examples.
Recognise that a published benchmark result does not establish fitness for a specific task.
L2 Applied
6 statementsApply defined acceptance criteria to judge whether an AI output meets a stated requirement.
Produce a set of test cases covering expected, edge and failure conditions for a defined task.
Document evaluation results so that another person can reproduce the judgement.
Check whether a change to a prompt, model or configuration has altered previously acceptable output.
Record an observed failure with sufficient context for it to be reproduced and investigated.
Report evaluation findings, distinguishing measured results from impressions.
L3 Proficient
8 statementsDesign fitness criteria for an AI system before development, derived from its intended use.
Design an evaluation set representative of production conditions and resistant to overfitting.
Evaluate agreement between human raters and diagnose the causes of disagreement.
Justify the choice of automated, model-graded or human evaluation for a given task.
Design regression testing that detects degradation across model and configuration changes.
Critique a published benchmark or vendor claim, identifying what it does and does not establish.
Design production monitoring that detects drift, degradation and emerging failure modes.
Review an incident involving AI output and identify the assurance gap that allowed it.
L4 Advanced
4 statementsEstablish evaluation standards and evidence requirements applying across an organisation’s AI systems.
Define the thresholds at which a system may enter, remain in, or must be withdrawn from use.
Commission independent evaluation where internal assurance is insufficient or conflicted.
Hold accountable those responsible for assurance evidence, including where evidence is absent.
Known gaps, open questions and contested points
Published because the record is more useful than the appearance of completeness.
Second domain drafted, and the domain that makes the framework credible to regulated buyers.
Most competing offerings omit it entirely, which is a significant differentiator and a significant responsibility.
L3 carries eight statements against six at L2. Proficient being the largest level is defensible for an assurance domain but is recorded here as a deliberate choice.