SETTING GLOBAL STANDARDS FOR TRUSTED AI CREDENTIALSAI Competence Framework v1.29 · current release
D6

Evaluation and assurance

PUBLISHED

Establishing with evidence whether a system is good enough, and when it stops being.

Statements
22 · L1 4 · L2 6 · L3 8 · L4 4
Version
v1.29
Last reviewed
30.08.2026
In scope
Defining fitness criteria in advance of building
Test case and evaluation set construction
Golden datasets and their maintenance
Human evaluation design, rater agreement and disagreement diagnosis
Automated and model-graded evaluation, and the limits of each
Regression testing across model, prompt and configuration changes
Benchmark interpretation, and what a benchmark result does not establish
Production monitoring, drift and degradation detection
Incident detection, investigation and post-incident review
Documentation of assurance evidence to a standard others can rely on
Out of scope
ExcludedWhere it lives
Formal audit engagement, independence and reportingO3
Research-depth statistical methodologyOut of framework
Setting the criteria themselvesD5
Boundary notes

Where this domain abuts another, and how the line is drawn

Against D5 Solution Design

D5 defines what good looks like. D6 establishes whether it was achieved.

Against overlay O3

D6 produces assurance evidence. O3 covers the professional obligations of the person auditing it.

Against D8 Security

D6 detects a security failure as a quality failure. D8 covers the adversary that caused it.

Statements by level

Every identifier is a permanent address. Indicators are normative; they state what would be observed in a person who meets the statement.

L1 Aware

4 statements
D6.L1.01Knowledge

Describe why AI output requires evaluation before it is relied upon.

Indicators
States that output can be fluent and confident while being wrong
Distinguishes plausibility from correctness
Identifies at least one consequence of relying on unevaluated output in their own work
D6.L1.02Knowledge

Distinguish a system that has been demonstrated from one that has been evaluated.

Indicators
States that a working example does not establish general performance
Recognises selection of favourable examples in a demonstration
Identifies what evidence would be needed to make a fitness claim
D6.L1.03Knowledge

State what a test case is and why evaluation requires more than a small number of examples.

Indicators
Describes a test case as an input paired with an expected outcome
States that non-determinism means a single pass is not a result
Recognises that easy cases alone do not establish fitness
D6.L1.04Judgement

Recognise that a published benchmark result does not establish fitness for a specific task.

Indicators
States that a benchmark measures performance on the benchmark
Identifies that the intended task may differ materially from what was measured
Escalates rather than relying on a vendor claim alone

L2 Applied

6 statements

Apply defined acceptance criteria to judge whether an AI output meets a stated requirement.

Indicators
Applies each criterion separately rather than forming a global impression
Records the criterion a failing output failed against
Refers cases the criteria do not cover rather than deciding them

Produce a set of test cases covering expected, edge and failure conditions for a defined task.

Indicators
Includes cases expected to fail, not only cases expected to pass
Covers realistic inputs rather than convenient ones
States the expected outcome for each case before running it

Document evaluation results so that another person can reproduce the judgement.

Indicators
Records system, version, configuration and date for each run
Records the criteria applied and their source
Retains the inputs and outputs, not only the verdict

Check whether a change to a prompt, model or configuration has altered previously acceptable output.

Indicators
Re-runs an existing case set after a change rather than testing only the change
Identifies output that changed without an intended cause
Reports regressions before the change is released

Record an observed failure with sufficient context for it to be reproduced and investigated.

Indicators
Captures the input, the output and the surrounding configuration
States what was expected and how the output differed
Distinguishes a reproducible failure from a single occurrence

Report evaluation findings, distinguishing measured results from impressions.

Indicators
Separates what was measured from what was inferred
States the size and composition of the evaluation set
States the limitations of the evaluation performed

L3 Proficient

8 statements

Design fitness criteria for an AI system before development, derived from its intended use.

Indicators
Derives criteria from the consequence of error in the intended use
Sets thresholds that are measurable rather than aspirational
Records the criteria and their basis before development begins

Design an evaluation set representative of production conditions and resistant to overfitting.

Indicators
Samples from realistic distributions rather than curated examples
Holds back cases not used during development
Identifies where the set is unrepresentative and states the consequence

Evaluate agreement between human raters and diagnose the causes of disagreement.

Indicators
Measures agreement rather than assuming it
Distinguishes ambiguous criteria from inconsistent application
Revises criteria or rater guidance in response to identified causes
D6.L3.04Judgement

Justify the choice of automated, model-graded or human evaluation for a given task.

Indicators
States what each method can and cannot establish for the task
Identifies where model-graded evaluation inherits the failure modes of the system evaluated
Records the reasoning so the choice can be reviewed

Design regression testing that detects degradation across model and configuration changes.

Indicators
Defines what constitutes a regression before changes occur
Covers behaviours that must not change, not only those that must work
Establishes when the test set itself requires revision
D6.L3.06Judgement

Critique a published benchmark or vendor claim, identifying what it does and does not establish.

Indicators
Identifies what was measured, on what data, under what conditions
States the gap between what was measured and the intended use
Distinguishes an unsupported claim from a supported claim about something else

Design production monitoring that detects drift, degradation and emerging failure modes.

Indicators
Selects signals that would move before harm occurs, not after
Sets thresholds that trigger action rather than notification alone
Provides for failure modes not anticipated at design time
D6.L3.08Judgement

Review an incident involving AI output and identify the assurance gap that allowed it.

Indicators
Traces the incident to the evaluation, monitoring or criterion that did not catch it
Distinguishes a gap in evidence from a gap in process
Produces a change to the assurance approach, not only to the system

L4 Advanced

4 statements
D6.L4.01Practice

Establish evaluation standards and evidence requirements applying across an organisation’s AI systems.

Indicators
Defines the minimum evidence required before any system may be relied upon
Sets requirements proportionate to consequence rather than uniform
Establishes how compliance with the standard is itself verified
D6.L4.02Judgement

Define the thresholds at which a system may enter, remain in, or must be withdrawn from use.

Indicators
States withdrawal conditions in advance, not at the point of failure
Assigns the authority to withdraw and the obligation to exercise it
Provides for withdrawal where evidence is absent, not only where it is negative
D6.L4.03Judgement

Commission independent evaluation where internal assurance is insufficient or conflicted.

Indicators
Identifies where internal evaluation cannot be relied upon and states why
Specifies scope and independence conditions for the commissioned work
Acts on findings that are unwelcome
D6.L4.04Practice

Hold accountable those responsible for assurance evidence, including where evidence is absent.

Indicators
Treats absence of evidence as a finding rather than a neutral state
Assigns named responsibility for assurance, not collective responsibility
Escalates where accountability is declined
Editorial notes

Known gaps, open questions and contested points

Published because the record is more useful than the appearance of completeness.

Second domain drafted, and the domain that makes the framework credible to regulated buyers.

Most competing offerings omit it entirely, which is a significant differentiator and a significant responsibility.

L3 carries eight statements against six at L2. Proficient being the largest level is defensible for an assurance domain but is recorded here as a deliberate choice.

This page displays version 1.29 · last reviewed 30.08.2026