Implementation Playbooks4 min read

AI Training for Employees: Run a Three-Task Baseline First

Before choosing AI training for employees, test three observable skills: selecting a suitable task, protecting inputs and verifying outputs.

Bokili Editorial· Verified August 25, 2026
ShareX
Three-task employee AI training baseline for task choice, safe inputs and output verification

AI training for employees should start with evidence of what people can do on a real-looking task—not a confidence survey or a generic course list. A short three-task baseline can reveal whether an employee can choose an appropriate use, protect the input and verify the output. L&D can then select training around observed gaps instead of sending everyone through the same material.

This is a practical training diagnostic, not a legal test and not an employee ranking. Run it with synthetic or approved content, score only observable behaviour and use the result to choose support. The goal is a fair starting point for learning.

Why baseline AI training for employees before choosing a course

The European Commission’s current AI-literacy Q&A says organisations should consider people’s technical knowledge, experience, education and training, as well as the AI systems and context involved. It also states that Article 4 does not oblige employers to measure every employee’s knowledge or guarantee a specific individual level. A task baseline is therefore a design choice: it helps make training relevant without being misrepresented as a compliance certificate.

NIST’s AI Use Taxonomy adds a useful task lens. It describes human-AI activities around goals and outcomes, independent of a particular technique or industry. That supports a baseline built around work behaviours rather than product trivia. OECD research likewise reports that skills remain important for effective generative-AI use and that training is not yet common among SMEs.

The three-task baseline

  1. 1

    1. Choose

    Give three fictional work requests. Ask the learner which one is suitable for an approved AI tool, which needs tighter boundaries and which should stay with a person.

  2. 2

    2. Protect

    Show a draft prompt containing one piece of unnecessary personal or confidential information. Ask the learner to remove or replace it and state the permitted source material.

  3. 3

    3. Verify

    Provide a short AI-generated answer with one unsupported claim. Ask the learner to trace it to the supplied sources, correct it and decide whether human review or escalation is needed.

Score four observable behaviours

1

Task fit

Identifies whether AI is appropriate for the stated goal and consequence.

2

Input boundary

Uses only approved, necessary information and spots what should be removed.

3

Output check

Tests important claims against the supplied evidence rather than trusting fluency.

4

Judgment

Knows when to correct, ask for review, escalate or stop.

Use a simple 0–2 scale for each behaviour: 0 means the action is absent or unsafe, 1 means it appears after a cue, and 2 means it is demonstrated independently with a clear reason. Keep the score descriptive. It should guide the next exercise, not become a performance rating. If employment decisions could be affected, involve HR, worker representatives and appropriate legal or equality expertise before using any assessment.

Worked example: a procurement summary

Reading is a start. Practice makes it stick.

Start learning

A fictional buyer receives a supplier note and must prepare a six-line comparison for a manager. In task one, the learner should recognise that summarising approved public specifications is suitable, while entering a confidential bid or asking AI to make the award decision is not. In task two, the learner replaces a real contact name and unpublished price with synthetic placeholders. In task three, the AI draft claims that one supplier is ‘fully compliant’, although the supplied documents only say that certification is pending. The learner should correct the claim, point to the source and flag the final decision for a person.

Observed resultTraining response
Chooses unsafe or vague usesTask-fit gapPractise use selection, consequences and stop conditions
Leaves sensitive detail in the promptInput-boundary gapPractise data minimisation with approved examples
Accepts the unsupported claimVerification gapPractise source tracing and evidence labels
Corrects the draft but makes the decisionJudgment gapPractise review ownership and escalation
Handles all three independentlyReady for transferMove to a fresh role-relevant task

Keep the baseline fair and useful

Seven safeguards for the diagnostic

  • Use the same instructions, time and approved tools for comparable learners.
  • Use synthetic, public or otherwise approved content.
  • Test ordinary work behaviours, not obscure product features.
  • Explain that the purpose is training selection, not surveillance.
  • Allow accessibility adjustments without changing the target behaviour.
  • Keep written anchors for 0, 1 and 2 on every behaviour.
  • Review patterns at cohort level before drawing conclusions about individuals.

Turn the result into a learning path

A baseline should lead directly to practice. Use the role-based AI skills matrix to name expected capability, the risk-based literacy guide to vary depth by context, the fair development-goal guide to set one next step and the four-week ChatGPT curriculum when that tool is in scope. Bokili’s HR and L&D page explains how short practice can support an organisation-wide learning plan.

Draft the baseline in ten minutes
  1. Choose one low-risk task that many employees recognise.
  2. Write one suitable AI use, one use needing tighter boundaries and one use that should remain human.
  3. Add one unnecessary sensitive detail to a synthetic prompt.
  4. Add one unsupported claim to a synthetic AI output and attach two short source notes.
  5. Write the four scoring anchors and the training action for each possible gap.

Do not build a hidden test

Tell employees what is being observed, how the result will be used and who can see it. A baseline works when it earns trust and directs practice. It fails when it becomes an unexplained score.

The best baseline is small enough to run before course selection and concrete enough to change that choice. Three tasks can show whether the first need is use selection, input safety, verification or human judgment. That is a stronger starting point for AI training for employees than assuming every learner has the same gap.

Sources

  1. AI Literacy — Questions & AnswersEuropean Commission
  2. AI Use Taxonomy: A Human-Centered ApproachNIST
  3. Are SMEs prepared for generative AI?OECD
  4. Bokili for HR and L&DBokili
ShareX

Reading is a start. Practice makes it stick.

Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.

Start learning

Keep reading