AI for Business4 min read

How to Measure AI Training Beyond Completion Rates

Use the MOVE evidence ladder to show whether AI learning changes practice, work quality and safe decision-making—not just course completion.

Bokili Editorial· Verified August 11, 2026
ShareX
A five-stage evidence ladder rises from course completion to safe workplace decisions.

A completion rate answers a useful but narrow question: who reached the end of the learning activity? It does not tell you whether people can frame a task, protect data, check an output, improve the work or recognise when not to use AI.

The measurement problem is not to find one perfect AI-adoption number. It is to build a short chain of evidence from practice to workplace behaviour to outcomes, while keeping safety and quality visible.

Measure the capability, not the click

NIST’s AI Risk Management Framework treats measurement as a mix of quantitative, qualitative or mixed methods used to analyse, assess, benchmark and monitor AI risks and impacts. Its Measure function calls for appropriate metrics, documented testing, regular reassessment and evidence that informs management decisions. That is a better model for learning teams than a dashboard built around attendance alone.

Activity metricCapability evidence
QuestionDid the learner finish?Can the learner perform the required behaviour?
UnitCourse, seat or moduleMission, decision or work sample
SignalExposurePractice, judgement and transfer
OwnerLearning teamLearning plus workflow and risk owners
ActionSend a reminderCoach, redesign, adjust controls or escalate

Use the MOVE evidence ladder

MOVE: four layers of training evidence

1

M — Missions attempted

Track whether people practise representative tasks, not only watch or read. Capture attempts, retries and where they ask for help.

2

O — Observable behaviour

Score a small set of visible actions: selecting an approved tool, protecting input, stating assumptions, checking sources, recording review and escalating.

3

V — Value and work quality

Sample whether the final work is clearer, more accurate, faster to revise or easier to audit. Use a baseline and a reviewer who understands the task.

4

E — Escalation judgement

Test whether people pause, refuse or route work correctly when data, uncertainty or consequence crosses a boundary.

Do not collapse MOVE into one vanity score

A high-use team can still produce weak or unsafe work. Keep adoption, quality and risk indicators visible as separate signals.

Define one decision before choosing metrics

Start with the management decision the evidence must support. “Prove the programme worked” is too vague. Better decisions include: which team needs coaching, which workflow is ready to scale, whether a control is understood, or where the learning task itself is unrealistic.

Reading is a start. Practice makes it stick.

Start learning

Build a minimal measurement design

  1. 1

    1. Name the behaviour

    Write one action an observer can see, such as opening the cited source before using a factual claim.

  2. 2

    2. Capture a baseline

    Give a small representative sample the task before training or record current work quality with a consistent rubric.

  3. 3

    3. Add a practice signal

    Use a realistic mission with an error, ambiguity or data-boundary decision. Log attempts and coaching needs.

  4. 4

    4. Sample transfer

    After an appropriate interval, review a small set of real or simulated work using the same behaviour rubric.

  5. 5

    5. Decide the response

    Set in advance what evidence will trigger coaching, workflow redesign, control changes or escalation.

Worked example: AI-assisted client brief

A consulting team learns to use an approved assistant for first drafts. Completion reaches 94%, but that alone cannot answer whether the workflow is ready to scale. The team selects four behaviours: remove restricted data, state the client question, verify every material external claim and record a human reviewer.

Evidence layerExample indicator
MissionsPercentage completing a realistic brief; retries before successWhere learners abandon or request coaching
Observable behaviourShare protecting inputs and opening decisive sourcesShare recording a named reviewer
Value and qualityRubric change in accuracy, clarity and revision effortTime saved only when quality is maintained
EscalationCorrect response to a planted unsupported claimCorrect routing of a sensitive-data scenario

The example deliberately combines numbers with review. NIST notes that measuring AI risk may require qualitative and quantitative methods and that metrics should provide meaningful information. A rubric comment can explain why a score moved; a count can show whether the pattern is widespread.

Avoid four measurement traps

Metric design check

  • Do not treat login frequency as proof of useful adoption.
  • Do not claim time savings without checking rework and quality.
  • Do not measure only easy tasks while scaling consequential ones.
  • Do not reward output volume when verification is required.
  • Do not compare teams with different tools, tasks or access as if contexts were equal.
  • Do not collect employee content or telemetry beyond a clear, lawful purpose.
  • Do not hide uncertainty in a single composite score.
  • Do document who reviews the evidence and what decision follows.
Upgrade one completion dashboard
  1. Choose one AI course with a completion metric.
  2. Write the workplace behaviour it is supposed to change.
  3. Design one five-minute mission that makes that behaviour visible.
  4. Choose one work-quality criterion and one safety boundary.
  5. Decide how you will sample transfer without collecting unnecessary content.
  6. Write the coaching or control decision each signal will trigger.

A good measurement system does not merely justify the learning programme. It helps the organisation decide where capability is real, where confidence exceeds skill and where the workflow or control—not the learner—needs to change.


Bokili supplies the mission layer in MOVE: short practice that makes behaviour visible through attempts, decisions and feedback. Pair that evidence with careful work sampling and accountable operational review.

Sources

  1. Artificial Intelligence Risk Management Framework (AI RMF 1.0)NIST
  2. AI RMF Core — MeasureNIST AI Resource Center
  3. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNIST
ShareX

Reading is a start. Practice makes it stick.

Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.

Start learning

Keep reading