SETTING GLOBAL STANDARDS FOR TRUSTED AI CREDENTIALSAI Competence Framework v1.29 · current release
Guidance note

Telling L2 from L3

Two tests that settle most level disputes: who chose the approach, and what happens when the situation is not the one the procedure anticipated.

The boundary between Applied and Proficient is where most mapping disputes end up, and almost none of them are really about difficulty. A hard task performed by following someone else’s procedure is L2. An easy task where the person decided what the approach should be, and could defend that decision, is L3.

Test one: who chose the approach

At L2 the approach is given. Acceptance criteria exist and the person applies them; a test-case template exists and the person fills it; the escalation route is defined and the person uses it. Competence at this level is doing the work correctly and recording it so that someone else can see what was done.

At L3 the person produces the approach. They derive fitness criteria from the intended use rather than receiving them, they decide that a task needs human evaluation rather than model-graded evaluation, and they can say why. If you remove the procedure and the work still happens correctly, you are looking at L3.

Test two: what happens at the edge of the procedure

Give the candidate a case the procedure does not cover. An L2 response refers it, and refers it well: it identifies that the criteria do not reach this case and escalates rather than guessing. That is not a deficiency; it is the correct L2 behaviour and one of the indicators.

An L3 response resolves it, and records the reasoning so the resolution can be reviewed. The distinction is not confidence. A confident guess is worse at both levels.

Difficulty is not level

Correction · 29.08.2026

An earlier version of this note used a complex regression-testing example to illustrate L3. That was a poor example: the complexity came from the tooling, not from the judgement, and two respondents were right to say the example would push assessors towards testing tool familiarity. The example below replaces it.

A junior analyst running a two-hundred-case evaluation set overnight is doing a large amount of L2 work. A lead who looks at a proposed evaluation set and says it is unrepresentative of production, states which distribution it misses, and accepts the delay that fixing it causes, is doing a small amount of L3 work. Volume, tooling and hours are not the variable.

What this means for a coverage map

If a course teaches learners to apply criteria that the course itself supplies, map it at L2 and say so. Mapping it at L3 because the subject matter is advanced is the single most common error in the coverage maps that reach the Registrar, and it is the one that most damages a learner: they hold a credential asserting they can design an approach, and their first real task is to design one.

Where you genuinely cannot tell, map lower and state the uncertainty in the exclusions. An under-claimed map is a smaller problem than an over-claimed one, for you and for the person holding the credential.

What this article discusses

Framework material referenced

These links run one way. The article points at the specification; the specification does not cite the article as guidance.

D6Evaluation and assurance
D6.L2.01Apply defined acceptance criteria to judge whether an AI output meets a stated requirement.
D6.L2.02Produce a set of test cases covering expected, edge and failure conditions for a defined task.
D6.L3.01Design fitness criteria for an AI system before development, derived from its intended use.
D6.L3.04Justify the choice of automated, model-graded or human evaluation for a given task.

Discusses framework version 0.1 · the article itself carries no version

Revision history

1 revision since publication

29.08.2026Second worked example rewritten after three consultation responses said the original conflated task difficulty with level. The paragraph on difficulty is new.
26.08.2026First published.

A notice is never edited. An article may be, and every substantive change appears above with the date it was made.

The author

Sofia Lindqvist

Editorial lead, D6 Evaluation and assurance

Leads the editorial group for D6 and drafted the twenty-two statements entered at version 0.1, with two technical reviewers per statement.

Works in assurance outside this organisation; that employment is declared below because a reader is entitled to know whose practice shaped the wording.

Declared interests, in full
Osei Assurance PartnersDeclared

Employed as lead assurance consultant. Osei Assurance Partners neither awards nor delivers credentials against this framework and holds no listing on the register.

AI Certification StandardsDeclared

Unpaid appointment to the editorial group. No economic interest.

Cite this article

Sofia Lindqvist (2026) “Telling L2 from L3”, Guidance note, AI Certification Standards. Non-normative. Available at aicertificationstandards.org/articles/telling-l2-from-l3 (accessed date).

Cite the article as an article. If you need to cite the competence itself, cite the statement: a credential specification or coverage map should reference statement identifiers, never this page.

Article last updated 29.08.2026 · page last reviewed 30.08.2026 · non-normative, not part of any framework version