Assessment design
The framework says what competence is. It deliberately does not prescribe how to assess it. This page is the practical consequence: which methods reach which statements, how to write items from indicators, and how to sample.
| Method | Knowledge | Skill | Judgement | Practice | Note |
|---|---|---|---|---|---|
| Selected response (multiple choice) | Reaches | No | No | No | Efficient for Knowledge. Using it beyond Knowledge is the single most common design failure. |
| Short written answer | Reaches | Partly | Partly | No | Reaches Skill only where the product of the skill is text. |
| Extended written case analysis | Reaches | Partly | Reaches | No | Strong for Judgement when the case has competing constraints and no single right answer. |
| Practical task with marked artefact | Reaches | Reaches | Partly | No | The default for Skill. |
| Simulation or lab exercise | Reaches | Reaches | Reaches | No | Can reach three types at once. Cannot show habitual practice. |
| Oral examination or defence | Reaches | Partly | Reaches | Partly | Excellent for Judgement, weak for reliability unless structured and double-marked. |
| Portfolio of real work with provenance | Reaches | Reaches | Reaches | Reaches | The only instrument that reaches Practice. |
| Workplace observation with attestation | Partly | Reaches | Reaches | Reaches | Reaches Practice directly. Dependent on the attester's own competence. |
Write from the indicators, not from the statement text
An item written from the statement text alone tends to test whether the candidate can restate it. An item written from an indicator tests whether they can do it.
Check an AI output against the source material it claims to rest on.
Indicator used: “Identifies claims in the output that the source does not support.”
The second item produces a marked artefact and cannot be answered from general awareness.
Design an evaluation approach for a task with no established benchmark.
Indicator used: “Justifies the choice of approach against at least one rejected alternative.”
Judgement items must have more than one defensible answer.
Apply baseline hardening to a deployed AI system to a documented procedure.
Indicator used: “Records which controls were applied and which could not be, with reasons.”
The selected-response item tests recognition. The practical reaches the Skill statement.
When the level is bigger than the examination
A level in a populated domain can hold eight statements with up to four indicators each. Sampling is legitimate; concealed sampling is not.
Sampling across statements is defensible; sampling inside one is how a statement becomes half-assessed while looking whole.
State how many statements are drawn, from which pool, and whether the draw is stratified.
An unstratified draw eventually produces a sitting that is all Knowledge.
Over a cycle every claimed statement should appear.
If a statement can be missed entirely, the map must say assessed-by-sampling.
Where multiple choice cannot reach
Judgement and Practice statements need instruments that cost more to run.
Cheapest instrument that reaches Skill honestly.
Guard: publish the rubric structure, and double-mark a sample every cycle.
The justification is what carries the Judgement claim.
Guard: if markers agree too easily, the scenario has a right answer and is not reaching Judgement.
The only route to Practice statements.
Guard: require an interview on the portfolio.
Strong evidence, entirely dependent on the attester.
Guard: state what competence the attester must hold.
Six checks against tool-specific items in a tool-neutral framework
Items drift toward the tool the course happens to teach on. The result is a credential that certifies familiarity with one product.
Could a competent candidate who has never used your course's platform pass this item?
Does the item name a product, model, vendor, menu, keyboard shortcut or interface element?
Would the correct answer change if the vendor shipped a release tomorrow?
Is the setting supplied as a described environment rather than as a specific tool, in every practical?
When a practical must run in a real tool, is a second tool offered, or is the tool incidental to marking?
Has anyone outside the course's teaching team reviewed the item bank for platform drift in the last cycle?