SETTING GLOBAL STANDARDS FOR TRUSTED AI CREDENTIALSAI Competence Framework v1.29 · current release
You are reading the current version of the framework, v1.29.Permanent address for this version

Critique a published benchmark or vendor claim, identifying what it does and does not establish.

Type Judgement · introduced in version 0.1

Performance indicators

Normative. These state what would be observed in a person who holds the statement.

Identifies what was measured, on what data, under what conditions
States the gap between what was measured and the intended use
Distinguishes an unsupported claim from a supported claim about something else
Evidence examples

Non-normative. Illustrative of evidence an awarding body might accept; not a required form.

A worked artefact produced in the course of normal duties, with the reasoning recorded at the time
Attestation by a competent supervisor against the indicators above, not against a general impression
Relationships

Assumes

Statements a candidate is taken to hold already. Never at a higher level than this one.

D6.L1.04

Recognise that a published benchmark result does not establish fitness for a specific task.

Assumed by

Derived inverse. Statements that take this one as given.

O3.L3.02

Evaluate whether evidence supports the claim being made.

Related

Cross-domain relationships, stated in both directions and typed in the content model.

D1.L3.01

Assess whether a class of task lies within the capability of current systems.

D1.L3.02

Compare candidate models against the stated requirements of a defined task.

D1.L3.03

Explain the mechanism behind a failure mode to a non-specialist audience.

D1.L3.04

Assess how context length, retrieval and tool access change what a system can do.

D1.L3.05

Distinguish limits inherent to the approach from limits of a particular implementation.

D1.L3.06

Advise on the capability, cost and latency trade-off in system selection.

D1.L3.07

Evaluate a vendor’s description of a system against its observable behaviour.

D1.L3.08

Maintain currency as capability changes, and revise advice already given.

D7.L3.01

Assess the impact of a proposed AI use on the people it affects.

D7.L3.02

Design fairness requirements for a system and the evidence they require.

D7.L3.03

Establish transparency and explanation requirements for a system.

D7.L3.04

Advise on a use that is permitted but should not proceed.

D7.L3.05

Design the route by which an affected person can challenge an AI-influenced decision.

D7.L3.06

Assess the labour and skill effects of an AI deployment.

D7.L3.07

Apply the obligations of a regulated or professional context to AI use.

D7.L3.08

Review an AI-related incident for the ethical failure, not only the technical one.

X-MOD.L3.04

Document a model so that others can judge its fitness.

Referenced by

Entries on the register of conformance claims whose coverage map cites this statement.

No entry on the register cites this statement. This block is populated from the coverage maps of register entries as they are listed.

Provenance
IntroducedVersion 0.1
Last modifiedNot modified since introduction
Statusactive · stable
Version displayedv1.29
Permanent URLaicertificationstandards.org/framework/statements/D6.L3.06
Cite this statement

AI Certification Standards (2026) AICF D6.L3.06, D6 Evaluation and assurance, version 1.29. Available at aicertificationstandards.org/framework/v1.29/statements/D6.L3.06 (accessed date).

Propose an amendment

Statements change through the published process, not by correspondence

An amendment to this statement, its indicators or its relationships is proposed through the contribution process. Every submission is answered on the record, and the reasoning for acceptance or rejection is published in the release record for the version that follows.

Propose an amendment to D6.L3.06

Writing rules and the controlled verb list that govern how this statement is worded are published in the methodology.

This page displays version 1.29 · last reviewed 30.08.2026