Why AI output has to be checked
AI output always sounds certain, even when it is wrong. Which errors to expect and how to check AI output without redoing all the work.
AI output has to be checked because an AI system does not distinguish between what it knows and what it considers likely: a wrong answer sounds just as certain as a right one. Language models can invent facts, misread sources or make arithmetic errors, and predictive models can go out of date. Checking does not have to mean redoing everything. It means that every output can be traced back to a source and that you catch the errors AI typically makes in a targeted way.
Why does AI get things wrong differently from people?
An employee who is unsure usually says so. "I think it is EUR 12,000, but let me check." A language model rarely gives that kind of signal of its own accord. It writes the most likely text, and a sentence that sounds certain is often more likely than a hesitant one.
On top of that:
- Errors are not random. AI makes mistakes in predictable places: arithmetic, conflicting sources, missing information, long documents.
- Errors are well packaged. An invented contract clause is written in correct legal language. A wrong total sits in a neat table.
- People get lazy. Someone who has received ten correct answers checks the eleventh less carefully. This is called automation bias, and it is the real risk.
Which five errors should you expect?
1. Inventing
The model fills in what is missing. Ask for the notice period in a contract that does not state one, and you may get a plausible period back. This is called hallucination, and it happens mainly when the model is under pressure to give an answer.
2. Misreading
The right document, the wrong line. An indexation percentage from an appendix that no longer applies, an amount excluding VAT read as including VAT, a date in American format.
3. Arithmetic
Language models do not add up like a calculator. With a handful of numbers it often goes well; with large volumes it goes wrong invisibly. That is why figures should come from a query or function, see what is tool calling.
4. The wrong source
The answer is correctly derived, but from an outdated or incorrect source. See what happens when AI uses the wrong business data.
5. Overconfident conclusions
The model sees a correlation and presents it as a cause. "Revenue fell because of the price increase", when the data only shows that both happened in the same month.
What checking is not
Checking does not mean an employee recalculates every output. Then the AI delivers nothing. Nor does it mean asking a second AI whether the first is right; it makes similar mistakes.
Good checking is layered: the system checks what it can check itself, and a person checks what only a person can judge.
How do you check AI output? Three layers
Layer 1: automatic, on every output
Checks software can do without human judgement:
- Is every figure in the source? An amount in the output must be traceable to the data the system retrieved.
- Do totals reconcile? The total of the reported invoice lines must match the total in the ledger.
- Is the format valid? Does the customer number exist, is the date plausible, does the percentage fall within a reasonable range?
- Is there a source? Output without a reference to records is not passed on.
Layer 2: sampling, periodically
A person checks a sample against the source. For example twenty findings a month, chosen at random, plus all findings above a threshold amount. Keep track of how many were correct. That figure is your quality measure, and it belongs in your AI monitoring.
Layer 3: human judgement, for every action with consequences
Anything that goes to a customer, moves money or changes master data is reviewed by a person before it happens. Not because the AI is likely to be wrong, but because the cost of an error is high there and the check is cheap. See human-in-the-loop AI explained.
Worked example: what checking costs and what it returns
Worked example: suppose an AI system reports 50 possible leaks between contract and invoice every month, with an average value of EUR 900. On review, 80 percent turn out to be valid and 20 percent not.
Without checking, pushing everything through: 40 valid corrections, EUR 36,000. But also 10 invalid corrections: customers who receive a supplementary invoice that is wrong. Each of those costs a credit note, an apology and goodwill.
With checking: every finding comes with its source attached: the contract clause, the invoice lines, the difference. An employee reviews one in three minutes on average. For 50 findings that is two and a half hours a month. The 10 invalid ones are removed before a customer sees them.
The figures are assumptions. The point is the ratio: checking costs little when the source is right there, and a lot when someone has to go looking. That is why traceability is the most important property of usable AI output.
Checking differs by type of output
Not every AI output needs the same check.
- Figures (amounts, totals, counts): check automatically against the source. Do not trust a figure that does not come from a calculation on the source.
- Extractions (a term from a contract, a date from an email): check a sample against the original document, plus everything with a lot of money attached.
- Summaries and explanations: check that every statement points to a source. A summary may be shorter than the source, not different.
- Predictions (likelihood of churn, expected revenue): check afterwards, by putting prediction and outcome side by side.
- Drafts (emails, invoice text): check before they go out, always by a person.
Making checking easy
The design of the output determines how quickly someone can check it. Good AI output for revenue control contains, for each finding:
- What was found, in one sentence.
- The evidence: the records that show it, with references.
- The calculation, if there is an amount.
- How confident the system is, and why.
- A proposed action.
With those five elements, checking is a matter of looking, not searching. How to assess that confidence level is covered in how you know whether an AI recommendation is reliable.
Where it goes wrong
The output is pushed through directly. An AI summary of a contract is pasted into the CRM as the source of truth. Three months later someone works from it without knowing it was a summary.
Only the outcome is shown. "Customer X underpays by EUR 2,300 a year" with no supporting evidence. The reviewer has to look everything up, so they either do not check or trust it blindly.
Sampling stops. In the first month everyone checks everything. By month six nobody does. That is exactly when something changes in the data or the model.
Checking without feedback. An invalid finding is dismissed, but nobody looks at why it was invalid. The same error comes back next month.
Checklist
- Can every output be traced to records in a source system?
- Do figures come from a calculation or query, not from the model itself?
- Are totals automatically reconciled with the ledger?
- Is there a fixed monthly sample, with a named reviewer?
- Is the number of valid findings tracked?
- Is every invalid finding fed back to its cause?
- Does a person review everything that goes to a customer or touches money?
In a revenue intelligence system
In a system for revenue control, verifiability is not an extra but the core. A leak alert that nobody can check does not get acted on. RiOS therefore shows, with every signal, the data behind it, how confident the system is, what it costs in euros and which fix is proposed, and it is designed so that a finding that does not meet the confidence threshold is not shown. How checking fits into the rest of the AI setup is explained in how AI works within Revenue Intelligence.
Frequently asked questions
Will AI models not simply get better, making checks unnecessary?
Models do get better, but errors caused by wrong or missing data remain. And the better the model, the greater the risk that people stop checking. Checking remains necessary; only its form can become lighter.
Can I have a second AI check what the first one does?
As a supplement it can be useful, for example to check that a summary contains all the amounts from the source. As the only check, no, because both models make similar mistakes.
How much should I check?
Everything that has consequences outside your own systems. On top of that, a fixed sample of the rest. If the sample is error-free for months, you can reduce it, but never drop it.
What if an error still slips through?
Then you want to know quickly what happened and which other outputs could be wrong in the same way. For that you need a step-by-step log.
More in this cluster
- How does AI work within Revenue Intelligence?Start here
- AI architecture for Revenue Intelligence
- What is anomaly detection?
- What is predictive analytics?
- Predictive AI vs generative AI
- What is an AI agent?
- What are AI agents in RevOps?
- What is tool calling?