Critique a published benchmark or vendor claim, identifying what it does and does not establish.
Type Judgement · introduced in version 0.1
Normative. These state what would be observed in a person who holds the statement.
Non-normative. Illustrative of evidence an awarding body might accept; not a required form.
Assumes
Statements a candidate is taken to hold already. Never at a higher level than this one.
Recognise that a published benchmark result does not establish fitness for a specific task.
Assumed by
Derived inverse. Statements that take this one as given.
Evaluate whether evidence supports the claim being made.
Related
Cross-domain relationships, stated in both directions and typed in the content model.
Assess whether a class of task lies within the capability of current systems.
Compare candidate models against the stated requirements of a defined task.
Explain the mechanism behind a failure mode to a non-specialist audience.
Assess how context length, retrieval and tool access change what a system can do.
Distinguish limits inherent to the approach from limits of a particular implementation.
Advise on the capability, cost and latency trade-off in system selection.
Evaluate a vendor’s description of a system against its observable behaviour.
Maintain currency as capability changes, and revise advice already given.
Assess the impact of a proposed AI use on the people it affects.
Design fairness requirements for a system and the evidence they require.
Establish transparency and explanation requirements for a system.
Advise on a use that is permitted but should not proceed.
Design the route by which an affected person can challenge an AI-influenced decision.
Assess the labour and skill effects of an AI deployment.
Apply the obligations of a regulated or professional context to AI use.
Review an AI-related incident for the ethical failure, not only the technical one.
Document a model so that others can judge its fitness.
Referenced by
Entries on the register of conformance claims whose coverage map cites this statement.
No entry on the register cites this statement. This block is populated from the coverage maps of register entries as they are listed.
| Introduced | Version 0.1 |
|---|---|
| Last modified | Not modified since introduction |
| Status | active · stable |
| Version displayed | v1.29 |
| Permanent URL | aicertificationstandards.org/framework/statements/D6.L3.06 |
AI Certification Standards (2026) AICF D6.L3.06, D6 Evaluation and assurance, version 1.29. Available at aicertificationstandards.org/framework/v1.29/statements/D6.L3.06 (accessed date).
Statements change through the published process, not by correspondence
An amendment to this statement, its indicators or its relationships is proposed through the contribution process. Every submission is answered on the record, and the reasoning for acceptance or rejection is published in the release record for the version that follows.
Propose an amendment to D6.L3.06Writing rules and the controlled verb list that govern how this statement is worded are published in the methodology.
This page displays version 1.29 · last reviewed 30.08.2026