SETTING GLOBAL STANDARDS FOR TRUSTED AI CREDENTIALSAI Competence Framework v1.29 · current release
You are reading the current version of the framework, v1.29.Permanent address for this version

Assessment design

The framework says what competence is. It deliberately does not prescribe how to assess it. This page is the practical consequence: which methods reach which statements, how to write items from indicators, and how to sample.

Method × statement type
MethodKnowledgeSkillJudgementPracticeNote
Selected response (multiple choice)ReachesNoNoNoEfficient for Knowledge. Using it beyond Knowledge is the single most common design failure.
Short written answerReachesPartlyPartlyNoReaches Skill only where the product of the skill is text.
Extended written case analysisReachesPartlyReachesNoStrong for Judgement when the case has competing constraints and no single right answer.
Practical task with marked artefactReachesReachesPartlyNoThe default for Skill.
Simulation or lab exerciseReachesReachesReachesNoCan reach three types at once. Cannot show habitual practice.
Oral examination or defenceReachesPartlyReachesPartlyExcellent for Judgement, weak for reliability unless structured and double-marked.
Portfolio of real work with provenanceReachesReachesReachesReachesThe only instrument that reaches Practice.
Workplace observation with attestationPartlyReachesReachesReachesReaches Practice directly. Dependent on the attester's own competence.
Writing items

Write from the indicators, not from the statement text

An item written from the statement text alone tends to test whether the candidate can restate it. An item written from an indicator tests whether they can do it.

D6.L2.03

Check an AI output against the source material it claims to rest on.

Indicator used: “Identifies claims in the output that the source does not support.”

✗ Why is it important to check AI outputs against sources? (500 words)
✓ Here is an output and its three cited sources. List every claim the sources do not support, and mark which source was misused.

The second item produces a marked artefact and cannot be answered from general awareness.

D6.L3.02

Design an evaluation approach for a task with no established benchmark.

Indicator used: “Justifies the choice of approach against at least one rejected alternative.”

✗ Describe the steps in designing an evaluation.
✓ Propose an approach for evaluating a summarisation feature, name one approach you rejected, and say what your approach cannot tell the client.

Judgement items must have more than one defensible answer.

D8.L2.05

Apply baseline hardening to a deployed AI system to a documented procedure.

Indicator used: “Records which controls were applied and which could not be, with reasons.”

✗ Which of the following is a hardening control? (a)…(d)
✓ In the supplied environment, apply the hardening procedure provided. Submit the completed record, including any control you could not apply and why.

The selected-response item tests recognition. The practical reaches the Skill statement.

Sampling

When the level is bigger than the examination

A level in a populated domain can hold eight statements with up to four indicators each. Sampling is legitimate; concealed sampling is not.

01Sample statements, never indicators within a claimed statement

Sampling across statements is defensible; sampling inside one is how a statement becomes half-assessed while looking whole.

02Publish the sampling rule, not just the syllabus

State how many statements are drawn, from which pool, and whether the draw is stratified.

03Stratify by statement type

An unstratified draw eventually produces a sitting that is all Knowledge.

04Rotate across sittings and record the rotation

Over a cycle every claimed statement should appear.

05Do not count sampled-out statements as assessed

If a statement can be missed entirely, the map must say assessed-by-sampling.

Practical and portfolio

Where multiple choice cannot reach

Judgement and Practice statements need instruments that cost more to run.

Practical with marked artefact

Cheapest instrument that reaches Skill honestly.

Guard: publish the rubric structure, and double-mark a sample every cycle.

Scenario with defended decision

The justification is what carries the Judgement claim.

Guard: if markers agree too easily, the scenario has a right answer and is not reaching Judgement.

Portfolio with provenance

The only route to Practice statements.

Guard: require an interview on the portfolio.

Observation with attester competence check

Strong evidence, entirely dependent on the attester.

Guard: state what competence the attester must hold.

Tool neutrality

Six checks against tool-specific items in a tool-neutral framework

Items drift toward the tool the course happens to teach on. The result is a credential that certifies familiarity with one product.

CHECK 1

Could a competent candidate who has never used your course's platform pass this item?

CHECK 2

Does the item name a product, model, vendor, menu, keyboard shortcut or interface element?

CHECK 3

Would the correct answer change if the vendor shipped a release tomorrow?

CHECK 4

Is the setting supplied as a described environment rather than as a specific tool, in every practical?

CHECK 5

When a practical must run in a real tool, is a second tool offered, or is the tool incidental to marking?

CHECK 6

Has anyone outside the course's teaching team reviewed the item bank for platform drift in the last cycle?