A Confidence Percentage Is Not Evidence
A model-written confidence percentage is not proof. Replace it with four evidence states that make sources, checks, gaps and boundaries visible.

Ask an AI system how confident it is and it may return a neat percentage. That number looks measurable. In ordinary workplace use, however, it is usually generated text—not a tested probability. Unless the system has been evaluated and calibrated for the specific task, model and operating conditions, “92% confident” does not tell a reviewer that an answer has a 92% chance of being right.
This matters because a number can end discussion. A fluent market brief, supplier summary or policy answer can appear more trustworthy when it carries a precise score, even though the score came from the same process that produced the claim. The safer move is simple: ignore unsupported confidence percentages and label the evidence behind each material claim.
Why a confidence percentage needs calibration
Calibration is a relationship between stated confidence and observed performance. If a system assigns 80% confidence to many answers, roughly 80% of those answers should be correct under the evaluated conditions. That relationship must be tested on representative tasks. It cannot be created merely by adding “give a confidence score” to a prompt.
Research shows why the distinction matters. A 2024 ACL study tested six prompting methods in question-answering and found that prompting could change calibration while still producing overconfidence on some instances. A 2026 open-access clinical benchmark tested 48 models on 300 gastroenterology questions and found poor self-estimation across that particular setting. The authors stress human oversight. The benchmark is domain-specific, so it is not proof that every model behaves identically on every task. It is a concrete warning against treating verbalised confidence as validated reliability.
| Confidence number | Evidence state | |
|---|---|---|
| What it describes | The model’s generated expression of certainty | What a reviewer can actually trace or check |
| How it earns trust | Only through task-specific calibration data | Through sources, independent checks and visible gaps |
| What action follows | Often unclear or based on an arbitrary threshold | Review, resolve, limit use or stop |
| What remains visible | A single score | The reason a claim can or cannot be used |
Replace the score with four evidence states
The SAFE evidence states
Source-backed
A named, accessible source directly supports the material claim. The reviewer has opened it and checked that the context matches.
Assured
A person or independent method has checked the claim against a second reliable source, calculation, dataset or approved record.
Flagged
The claim may be plausible, but a source is missing, ambiguous, outdated or contradictory. It must not pass as established fact.
Excluded
The question sits outside the available evidence, approved data or reviewer’s competence. The workflow stops or routes to an owner.
SAFE is not a new model score. It is a work-status language. A claim can move from Flagged to Source-backed after a source is opened, or from Source-backed to Assured after an independent check. Excluded is not failure; it makes a boundary visible before someone fills the gap with a guess.
Worked example: the confident supplier brief
Reading is a start. Practice makes it stick.
Start learningImagine an operations team asks an assistant to compare two suppliers using public information. The draft says Supplier A has the stronger delivery record and reports “92% confidence”. The reviewer does not debate whether 92 should be 85. She removes the number and marks the underlying claims.
Evidence review of the draft
- Source-backed: each supplier’s published service area and stated delivery target.
- Assured: current pricing checked against the signed quotation and the finance record.
- Flagged: a claim that Supplier A has fewer delays, because the draft cites no comparable dataset.
- Excluded: a recommendation about contractual risk, which requires procurement and legal review.
- Decision: the brief may inform a shortlist only after the flagged claim is removed or supported.
The result is more useful than a confidence score. It shows which claims are ready, which need work and which belong to another owner. It also creates a review trail that a colleague can understand without reproducing the entire conversation.
Do not turn SAFE into another hidden score
The four states should not be collapsed into one percentage or used to reward people for producing more green labels. NIST’s Generative AI Profile treats trustworthiness as something organisations manage across design, use and evaluation. In practice, that means preserving context: task consequence, source quality, data limits, reviewer competence and the cost of an error.
Manager guardrails
- Define which claims are material enough to label.
- Require the source to be opened, not merely listed.
- Use a second check when an error could change money, rights, safety or reputation.
- Name the owner who can resolve a Flagged or Excluded item.
- Keep the original evidence with the final work product.
- Never use a model-written confidence percentage as the sole approval rule.
Connect the method to existing checks
SAFE complements Bokili’s guides on verifying AI output, setting a stop rule, running a workflow risk clinic and handing off AI-assisted work. Use the evidence state inside those broader workflows rather than creating a parallel approval system.
- Choose one low-risk AI-generated brief that contains three or four factual claims.
- Delete any self-reported confidence number.
- Open the cited material and label each claim Source-backed, Assured, Flagged or Excluded.
- Write the action required for every Flagged or Excluded item.
- Ask a colleague whether the labels make the next decision clearer.
- Save the labels with the reviewed brief, not only in the chat.
Precision is valuable when it comes from evidence and evaluation. A percentage without calibration only imitates that precision. Give reviewers a status they can inspect, a source they can open and a boundary they can act on.
Sources
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile — NIST
- Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language Models — Association for Computational Linguistics
- Across generations, sizes, and types, large language models poorly report self-confidence in gastroenterology clinical reasoning tasks — npj Gut and Liver
Reading is a start. Practice makes it stick.
Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.
Start learningKeep reading

AI Course for Beginners: Build One Safe Work Sample
Choose one low-consequence task, protect the inputs, define a quality bar and build a verified first AI work sample.

Separate Generation From Decision: A Two-Pass AI Template
Use AI to expand and challenge options, then make and record the accountable human choice in a separate pass.

AI Training for Employees on Shifts: A Frontline Playbook
Design AI training for employees in retail, operations and field roles with short practice, safe examples, fast feedback and next-shift transfer.