How do you know whether an AI recommendation is reliable?
An AI recommendation is reliable when the evidence holds, the confidence is justified and the system is demonstrably right more often than not.
An AI recommendation is reliable when you can verify the evidence in your own systems, when the system states how confident it is and why, and when you can show over a longer period that high-confidence recommendations were right more often than low-confidence ones. A confidence score on its own says little; only when you put it next to actual outcomes do you know whether you can use it. Without a source and without a measurable hit rate, a recommendation is an opinion with a number attached.
Three questions for every recommendation
Before you look at scores, there are three questions you can ask of any recommendation.
1. Can I see the evidence? A recommendation to send a customer a corrective invoice must point to the contract clause, the invoice lines and the calculated difference. If you can check those in a minute, the recommendation is verifiable. If you cannot, the system is asking for blind trust.
2. Does the reasoning hold? Does the recommendation follow logically from the evidence? "The customer pays less than last year, so there is a pricing error" does not hold if the customer also bought less. A good recommendation rules out alternative explanations or names them.
3. How confident is the system, and is that confidence justified? That is what the rest of this article is about.
What is a confidence score in AI?
Many AI systems give a confidence score with an output: a number between 0 and 1, a percentage, or a label such as high, medium or low. What that number means differs by type of AI.
In predictive models
A predictive model gives a probability. "This customer has a 0.78 probability of not renewing." If the model is well calibrated, that means: of all customers scored 0.78, roughly 78 percent do not renew. That is testable.
In rules and comparisons
A finding that comes from a hard comparison, contract price EUR 95 and invoiced EUR 85, is in principle certain, provided the source data is right. The uncertainty then lies not in the calculation but in the data: is this the right contract version, is there perhaps an agreement that is not in the system?
In language models
A language model that extracts an agreement from a contract has no reliable built-in confidence score. If you ask the model how confident it is, you get a number, but that number is also generated text. It tells you something, but not enough to rely on. Better signals are: did the model find a verbatim passage supporting the agreement, did several independent extractions give the same answer, does the format match what you expect?
Calibration: the only thing that counts
A confidence score is only useful if it is calibrated: if 90 percent confidence really is right about nine times out of ten. Many systems are overconfident. They report 90 percent where the reality is 65 percent.
You can test this yourself, with data you already have. For each recommendation, record the score it had and whether it proved right on review. After a few months, build a table:
| Confidence according to system | Number of recommendations | Correct on review | Actual hit rate |
|---|---|---|---|
| 90 to 100% | 120 | 110 | 92% |
| 70 to 90% | 80 | 58 | 73% |
| 50 to 70% | 60 | 27 | 45% |
In this example the top two groups are reasonably calibrated; the lowest group is too optimistic. Recommendations below 70 percent therefore deserve extra attention, or should not be shown at all.
The figures in the table are an example. The method is what matters: without this table you do not know what your system's scores are worth.
Worked example: choosing a threshold
Worked example: suppose a system finds 200 possible leaks a month with an average value of EUR 600. Review by an employee takes five minutes each. An unjustified action towards a customer costs you an estimated EUR 250 in time and goodwill.
With the calibration table above:
- Show everything: 200 reviews, 16.7 hours of work. All valid leaks are found, provided the reviewer stays sharp.
- Show only above 70 percent: with the distribution from the table, that is roughly 154 reviews, 12.8 hours. You then miss the valid leaks in the lowest group: about 46 recommendations, of which 45 percent are valid. That is around 21 leaks, at EUR 600 on average roughly EUR 12,500 a month.
- Lowest group separately, in a monthly bulk review: you keep the return but concentrate the work.
Which threshold is right depends on the balance between the value of a leak found, the cost of review and the cost of an error. That balance is different for every business. The principle that uncertain findings are not published is called fail-closed. More on what leaks cost is in how to calculate revenue leakage.
Signs that a recommendation is unreliable
- No source. An amount or conclusion without a reference to records.
- Always high confidence. A system that gives 95 percent on every recommendation does not discriminate.
- Source and conclusion diverge. The contract says 2.5 percent, the recommendation calculates with 3 percent.
- Inconsistent answers. The same situation produces a different recommendation when repeated.
- Outdated data. The recommendation is based on data from before a known change.
- Too good to be true. A leak of EUR 200,000 at a customer with EUR 50,000 annual revenue. Probably an integration error.
That last category often arises from the wrong business data, not from the model.
Reliability changes over time
A system that was well calibrated last quarter need not be now. Three things change it:
- Your data changes. A new pricing model, a migration to a different accounting package, a new product line. Patterns the system relied on no longer apply.
- The model changes. Language model providers release new versions. Behaviour can shift, even if nobody on your side changes anything.
- Your processes change. If your team systematically closes certain leaks, the remaining findings become rarer and often harder.
That is why the calibration table is not a one-off test but a fixed part of AI monitoring.
Checklist: assessing an AI system for reliability
Ask a vendor or your own team:
- Does every recommendation show the evidence from the source systems?
- Is there a confidence level per recommendation, and how is it determined?
- Is that level calibrated, and can you see the data?
- What happens to recommendations below a threshold? Are they shown, set aside or left out?
- Is it tracked how often recommendations were right?
- Are rejected recommendations used to improve the system?
- How does the system handle conflicting sources? Does it pick one, or report the conflict?
Reliability is a property of the system, not the model
A better language model does not automatically make recommendations more reliable. Reliability comes from the whole: correct source data, hard calculations in software, a justified confidence level, a threshold for publication and a person who reviews the rest. That whole is described in how AI works within Revenue Intelligence. The role of the person in it is worked out in human-in-the-loop AI explained.
In RiOS, every signal shows the data behind it, a confidence level, the impact in euros and a proposed step, and the platform is built on fail-closed publication: a finding that does not meet the confidence threshold does not reach the dashboard. How that works is shown on the page about the system.
Frequently asked questions
What is a good confidence score?
There is no universal cut-off. A score is good if it is calibrated, meaning that 80 percent really is right about eight times out of ten. Which threshold you use depends on what a missed finding and an unjustified action cost you.
Can I ask a language model how confident it is?
You can, but the answer is itself generated text and not reliably calibrated. Look instead at whether the model can point to a verbatim source and whether the answer matches other sources.
How long does it take to know whether a system is reliable?
You need enough reviewed recommendations to say something per confidence group. With dozens of recommendations a month, you have a first picture after a few months. Keep tracking it after that, because data and models change.
Is a recommendation with 100 percent confidence always right?
No. A hard comparison can be 100 percent certain about the calculation and still be wrong because the source data is incorrect or an agreement was made outside the system.
More in this cluster
- How does AI work within Revenue Intelligence?Start here
- AI architecture for Revenue Intelligence
- What is anomaly detection?
- What is predictive analytics?
- Predictive AI vs generative AI
- What is an AI agent?
- What are AI agents in RevOps?
- What is tool calling?