How to Measure AI Training Beyond Completion Rates
Use the MOVE evidence ladder to show whether AI learning changes practice, work quality and safe decision-making—not just course completion.

A completion rate answers a useful but narrow question: who reached the end of the learning activity? It does not tell you whether people can frame a task, protect data, check an output, improve the work or recognise when not to use AI.
The measurement problem is not to find one perfect AI-adoption number. It is to build a short chain of evidence from practice to workplace behaviour to outcomes, while keeping safety and quality visible.
Measure the capability, not the click
NIST’s AI Risk Management Framework treats measurement as a mix of quantitative, qualitative or mixed methods used to analyse, assess, benchmark and monitor AI risks and impacts. Its Measure function calls for appropriate metrics, documented testing, regular reassessment and evidence that informs management decisions. That is a better model for learning teams than a dashboard built around attendance alone.
| Activity metric | Capability evidence | |
|---|---|---|
| Question | Did the learner finish? | Can the learner perform the required behaviour? |
| Unit | Course, seat or module | Mission, decision or work sample |
| Signal | Exposure | Practice, judgement and transfer |
| Owner | Learning team | Learning plus workflow and risk owners |
| Action | Send a reminder | Coach, redesign, adjust controls or escalate |
Use the MOVE evidence ladder
MOVE: four layers of training evidence
M — Missions attempted
Track whether people practise representative tasks, not only watch or read. Capture attempts, retries and where they ask for help.
O — Observable behaviour
Score a small set of visible actions: selecting an approved tool, protecting input, stating assumptions, checking sources, recording review and escalating.
V — Value and work quality
Sample whether the final work is clearer, more accurate, faster to revise or easier to audit. Use a baseline and a reviewer who understands the task.
E — Escalation judgement
Test whether people pause, refuse or route work correctly when data, uncertainty or consequence crosses a boundary.
Do not collapse MOVE into one vanity score
A high-use team can still produce weak or unsafe work. Keep adoption, quality and risk indicators visible as separate signals.
Define one decision before choosing metrics
Start with the management decision the evidence must support. “Prove the programme worked” is too vague. Better decisions include: which team needs coaching, which workflow is ready to scale, whether a control is understood, or where the learning task itself is unrealistic.
Reading is a start. Practice makes it stick.
Start learningBuild a minimal measurement design
- 1
1. Name the behaviour
Write one action an observer can see, such as opening the cited source before using a factual claim.
- 2
2. Capture a baseline
Give a small representative sample the task before training or record current work quality with a consistent rubric.
- 3
3. Add a practice signal
Use a realistic mission with an error, ambiguity or data-boundary decision. Log attempts and coaching needs.
- 4
4. Sample transfer
After an appropriate interval, review a small set of real or simulated work using the same behaviour rubric.
- 5
5. Decide the response
Set in advance what evidence will trigger coaching, workflow redesign, control changes or escalation.
Worked example: AI-assisted client brief
A consulting team learns to use an approved assistant for first drafts. Completion reaches 94%, but that alone cannot answer whether the workflow is ready to scale. The team selects four behaviours: remove restricted data, state the client question, verify every material external claim and record a human reviewer.
| Evidence layer | Example indicator | |
|---|---|---|
| Missions | Percentage completing a realistic brief; retries before success | Where learners abandon or request coaching |
| Observable behaviour | Share protecting inputs and opening decisive sources | Share recording a named reviewer |
| Value and quality | Rubric change in accuracy, clarity and revision effort | Time saved only when quality is maintained |
| Escalation | Correct response to a planted unsupported claim | Correct routing of a sensitive-data scenario |
The example deliberately combines numbers with review. NIST notes that measuring AI risk may require qualitative and quantitative methods and that metrics should provide meaningful information. A rubric comment can explain why a score moved; a count can show whether the pattern is widespread.
Avoid four measurement traps
Metric design check
- Do not treat login frequency as proof of useful adoption.
- Do not claim time savings without checking rework and quality.
- Do not measure only easy tasks while scaling consequential ones.
- Do not reward output volume when verification is required.
- Do not compare teams with different tools, tasks or access as if contexts were equal.
- Do not collect employee content or telemetry beyond a clear, lawful purpose.
- Do not hide uncertainty in a single composite score.
- Do document who reviews the evidence and what decision follows.
- Choose one AI course with a completion metric.
- Write the workplace behaviour it is supposed to change.
- Design one five-minute mission that makes that behaviour visible.
- Choose one work-quality criterion and one safety boundary.
- Decide how you will sample transfer without collecting unnecessary content.
- Write the coaching or control decision each signal will trigger.
A good measurement system does not merely justify the learning programme. It helps the organisation decide where capability is real, where confidence exceeds skill and where the workflow or control—not the learner—needs to change.
Bokili supplies the mission layer in MOVE: short practice that makes behaviour visible through attempts, decisions and feedback. Pair that evidence with careful work sampling and accountable operational review.
Sources
Reading is a start. Practice makes it stick.
Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.
Start learningKeep reading

AI Course for Beginners: Build One Safe Work Sample
Choose one low-consequence task, protect the inputs, define a quality bar and build a verified first AI work sample.

Separate Generation From Decision: A Two-Pass AI Template
Use AI to expand and challenge options, then make and record the accountable human choice in a separate pass.

AI Training for Employees on Shifts: A Frontline Playbook
Design AI training for employees in retail, operations and field roles with short practice, safe examples, fast feedback and next-shift transfer.