Before You Scale an AI Pilot, Write Its Exit Rules
Define the evidence for repeating, scaling or stopping an AI pilot before a persuasive demo makes the decision for you.

An AI pilot should not end with a presentation called “promising results”. It should end with a decision. Before the first participant, task or test set enters the pilot, write the evidence that would justify one of three exits: repeat the experiment, scale the workflow, or stop and redirect the effort.
Without those rules, a convincing demo can be mistaken for operational readiness. The UK Government AI Playbook describes phased experiments, accuracy thresholds and evaluation across the project life cycle. NIST likewise recommends documenting test sets, metrics, methods and performance outcomes so that evaluation is repeatable. The practical implication is simple: define the decision before seeing the result.
Three exit rules for an AI pilot
Repeat, scale or stop
Repeat to answer one gap
Run another bounded cycle only when value remains plausible and one named uncertainty can be tested with a specific change.
Scale with controls
Expand only when task quality, operational safeguards and ongoing support all meet their pre-agreed thresholds.
Stop or redirect
End the pilot when the outcome is not useful, the risk cannot be controlled, or the review and support burden defeats the value.
These are not three moods. Each needs evidence. “People liked it” may support continued exploration, but it cannot prove that the workflow is accurate, safe or supportable at larger volume. “The model produced a good example” says little about how often a reviewer must intervene, which cases fail, or whether the team can maintain the process.
Measure readiness, not enthusiasm alone
| Demo evidence | Scale evidence | |
|---|---|---|
| Task result | One persuasive output | A defined test set assessed against explicit acceptance criteria |
| Human work | A reviewer corrected the example | Review time, escalation points and decision rights are understood |
| Risk | No problem appeared in the demo | Known failure modes have controls, owners and stop conditions |
| Operations | The pilot team made it work | A normal team can repeat, monitor and support the workflow |
Use three evidence gates. The task gate asks whether the output is good enough for its intended purpose across representative cases. The control gate asks whether data, approval, escalation and human review rules actually work. The support gate asks whether the workflow can be repeated without relying on the pilot’s most enthusiastic expert.
A pilot can pass the task gate and still fail the support gate. That is not a disappointing footnote; it is the result. The GOV.UK Chat case study, for example, found that highly manual quality assurance would not scale and identified the need for a quality-assessed question set. The lesson is broader than chatbots: scaling changes the operating problem.
Worked example: a supplier-onboarding brief
Reading is a start. Practice makes it stick.
Start learningImagine a procurement team pilots an AI workflow that turns fictional supplier documents into an onboarding brief. The team uses a fixed pack of sample documents, including a missing certificate, an inconsistent payment term and one clear case. A human reviewer owns the final decision.
Write the exit card before the test
- 1
Define the task threshold
The brief must preserve every material fact, flag missing evidence and avoid inventing a conclusion. The test set and scoring method are recorded.
- 2
Define the control threshold
No live supplier data enters the exercise. Every exception has an escalation route, and the reviewer knows which fields cannot be approved from AI output.
- 3
Define the support threshold
A second team member can run the workflow from the written instructions, diagnose common failures and record changes without help from the pilot designer.
- 4
Assign the decision
The sponsor records which failed threshold triggers a repeat, which combination permits limited scale, and which condition ends the pilot.
Suppose the briefs are accurate on straightforward cases but miss the inconsistent payment term. The correct exit is not “scale carefully”. It is a repeat decision with one named question: can the workflow reliably surface contractual contradictions after the prompt, source pack or review checklist is changed? The next cycle exists to answer that question, not to keep the pilot alive.
Prevent the endless pilot
A repeat rule needs its own limit. Specify the change, the evidence expected and the maximum number of cycles. If each round introduces a different objective, the organisation is no longer testing a pilot; it is postponing a choice. Preserve the change log so that improved results can be traced to a changed prompt, model, task boundary or review process.
Scale should also be staged. A workflow that passes within one team may need another controlled gate before it crosses functions, countries or risk levels. New users and operating conditions can expose failures absent from the first test. NIST notes that context affects which evaluation methods are appropriate and that measurements should be reassessed as operational settings change.
- Name the exact work task and the decision the pilot is meant to inform.
- Write one observable pass condition for task quality, controls and supportability.
- Define the single uncertainty that would justify a repeat cycle.
- Name one condition that forces a stop or redesign.
- Assign the person who records the decision and the evidence behind it.
Make the decision the deliverable
The pilot is not successful because it ran or because participants were engaged. It is successful when it produces trustworthy evidence for a decision. Define repeat, scale and stop before the work begins; then let the evidence choose the route.
Bokili helps teams practise the behaviours behind those gates in short, focused missions. Connect the exit card to related guidance on training the workflow constraint, turning dashboard signals into support and deciding what saved time becomes.
Sources
- Artificial Intelligence Playbook for the UK Government — UK Government
- AI RMF Playbook — Measure — NIST
- Planning & Evaluating — U.S. Office of Personnel Management
Reading is a start. Practice makes it stick.
Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.
Start learningKeep reading

ChatGPT Training for Employees: Build a Safe Practice Pack
Create realistic ChatGPT training for employees with fictional facts, edge cases, review criteria and an answer key—without importing live work data.

AI Learning Path: Pair Every Creation Skill With a Check
Build an AI learning path that pairs useful output with source, constraint, assumption and decision-boundary checks.

Use Gemini Notebook to Test One Claim Against One Source
A focused source check turns a fluent answer into inspectable evidence before you widen the notebook.