Responsible AI Practice4 min read

A Confidence Percentage Is Not Evidence

A model-written confidence percentage is not proof. Replace it with four evidence states that make sources, checks, gaps and boundaries visible.

Bokili Editorial· Verified August 22, 2026
ShareX
Reviewer replaces an AI confidence score with four evidence states tied to source material

Ask an AI system how confident it is and it may return a neat percentage. That number looks measurable. In ordinary workplace use, however, it is usually generated text—not a tested probability. Unless the system has been evaluated and calibrated for the specific task, model and operating conditions, “92% confident” does not tell a reviewer that an answer has a 92% chance of being right.

This matters because a number can end discussion. A fluent market brief, supplier summary or policy answer can appear more trustworthy when it carries a precise score, even though the score came from the same process that produced the claim. The safer move is simple: ignore unsupported confidence percentages and label the evidence behind each material claim.

Why a confidence percentage needs calibration

Calibration is a relationship between stated confidence and observed performance. If a system assigns 80% confidence to many answers, roughly 80% of those answers should be correct under the evaluated conditions. That relationship must be tested on representative tasks. It cannot be created merely by adding “give a confidence score” to a prompt.

Research shows why the distinction matters. A 2024 ACL study tested six prompting methods in question-answering and found that prompting could change calibration while still producing overconfidence on some instances. A 2026 open-access clinical benchmark tested 48 models on 300 gastroenterology questions and found poor self-estimation across that particular setting. The authors stress human oversight. The benchmark is domain-specific, so it is not proof that every model behaves identically on every task. It is a concrete warning against treating verbalised confidence as validated reliability.

Confidence numberEvidence state
What it describesThe model’s generated expression of certaintyWhat a reviewer can actually trace or check
How it earns trustOnly through task-specific calibration dataThrough sources, independent checks and visible gaps
What action followsOften unclear or based on an arbitrary thresholdReview, resolve, limit use or stop
What remains visibleA single scoreThe reason a claim can or cannot be used

Replace the score with four evidence states

The SAFE evidence states

1

Source-backed

A named, accessible source directly supports the material claim. The reviewer has opened it and checked that the context matches.

2

Assured

A person or independent method has checked the claim against a second reliable source, calculation, dataset or approved record.

3

Flagged

The claim may be plausible, but a source is missing, ambiguous, outdated or contradictory. It must not pass as established fact.

4

Excluded

The question sits outside the available evidence, approved data or reviewer’s competence. The workflow stops or routes to an owner.

SAFE is not a new model score. It is a work-status language. A claim can move from Flagged to Source-backed after a source is opened, or from Source-backed to Assured after an independent check. Excluded is not failure; it makes a boundary visible before someone fills the gap with a guess.

Worked example: the confident supplier brief

Reading is a start. Practice makes it stick.

Start learning

Imagine an operations team asks an assistant to compare two suppliers using public information. The draft says Supplier A has the stronger delivery record and reports “92% confidence”. The reviewer does not debate whether 92 should be 85. She removes the number and marks the underlying claims.

Evidence review of the draft

  • Source-backed: each supplier’s published service area and stated delivery target.
  • Assured: current pricing checked against the signed quotation and the finance record.
  • Flagged: a claim that Supplier A has fewer delays, because the draft cites no comparable dataset.
  • Excluded: a recommendation about contractual risk, which requires procurement and legal review.
  • Decision: the brief may inform a shortlist only after the flagged claim is removed or supported.

The result is more useful than a confidence score. It shows which claims are ready, which need work and which belong to another owner. It also creates a review trail that a colleague can understand without reproducing the entire conversation.

Do not turn SAFE into another hidden score

The four states should not be collapsed into one percentage or used to reward people for producing more green labels. NIST’s Generative AI Profile treats trustworthiness as something organisations manage across design, use and evaluation. In practice, that means preserving context: task consequence, source quality, data limits, reviewer competence and the cost of an error.

Manager guardrails

  • Define which claims are material enough to label.
  • Require the source to be opened, not merely listed.
  • Use a second check when an error could change money, rights, safety or reputation.
  • Name the owner who can resolve a Flagged or Excluded item.
  • Keep the original evidence with the final work product.
  • Never use a model-written confidence percentage as the sole approval rule.

Connect the method to existing checks

SAFE complements Bokili’s guides on verifying AI output, setting a stop rule, running a workflow risk clinic and handing off AI-assisted work. Use the evidence state inside those broader workflows rather than creating a parallel approval system.

Replace one confidence score in ten minutes
  1. Choose one low-risk AI-generated brief that contains three or four factual claims.
  2. Delete any self-reported confidence number.
  3. Open the cited material and label each claim Source-backed, Assured, Flagged or Excluded.
  4. Write the action required for every Flagged or Excluded item.
  5. Ask a colleague whether the labels make the next decision clearer.
  6. Save the labels with the reviewed brief, not only in the chat.

Precision is valuable when it comes from evidence and evaluation. A percentage without calibration only imitates that precision. Give reviewers a status they can inspect, a source they can open and a boundary they can act on.

Sources

  1. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNIST
  2. Fact-and-Reflection (FaR) Improves Confidence Calibration of Large Language ModelsAssociation for Computational Linguistics
  3. Across generations, sizes, and types, large language models poorly report self-confidence in gastroenterology clinical reasoning tasksnpj Gut and Liver
ShareX

Reading is a start. Practice makes it stick.

Bokili turns skills like this into ten-minute missions for your whole team, with instant feedback and progress you can see.

Start learning

Keep reading