AI monitoring explained
AI monitoring tracks what an AI system reads, decides and does, and whether it stays right. What to monitor, why agents need it and how to start.
AI monitoring is the continuous tracking of what an AI system does and whether it is still doing it well: which data it read, which actions it carried out, how often its outputs were correct, what it cost and when its behaviour changes. For AI that only advises, it is a quality check. For AI agents that act inside your systems, it is the only way to know what happened and what you need to put right when something goes wrong.
Why AI needs different monitoring from ordinary software
Ordinary software does the same thing every time. If an invoicing run works today, it will work tomorrow unless someone changes something. Monitoring is then mainly about whether the system is running.
AI behaves differently:
- The same question can produce a different answer. Language models have an element of variation in their output.
- The model can change without you doing anything. Providers release new versions, and behaviour can shift even within a version.
- The data changes. A predictive model trained on last year can perform worse this year. This is called drift.
- Errors are silent. An AI system does not crash when it produces a wrong finding. It simply gives an answer.
So monitoring AI is not just about whether it runs, but above all about whether it is still right.
Why AI agents without monitoring can be dangerous
With an AI agent there is an extra layer. An agent acts: it looks things up, updates records, prepares and sends. Without monitoring, three risks reinforce each other.
You do not know what it did. If an agent spends a week updating CRM records and you then discover an error, you need to know which records it touched, when, and on what basis. Without a log, all you know is that something is wrong.
You do not know why. An agent that has been misled by text in an email or a note behaves differently from what was intended. You can only trace that if you record what it read just before it acted.
It keeps going. A person who sees something odd stops. An agent without limits and without alerts makes the same mistake a hundred times.
Revoking a key stops what an agent will do next. Monitoring tells you what it has already done.
What do you monitor? Four layers
1. Activity
Record for every step:
- Which tool was called, with which parameters.
- Which data was viewed.
- What was changed, created or sent.
- With which key, under which account.
- The time.
This is the log. It is the foundation for everything that follows. How tool calls work is covered in what is tool calling.
2. Quality
Is what the AI says correct? For revenue control this can be measured concretely:
- Valid findings. Of all reported discrepancies, how many turn out on review to be a real leak?
- Missed findings. How many leaks were later found another way that the AI should have seen?
- Deviation from the source. Do the amounts in the output match the source?
- Predictions against reality. For predictive analytics: how often was the prediction right?
3. Behaviour
Is the system changing? Signals include:
- A sudden rise or fall in the number of findings.
- An agent that needs more steps for the same task.
- More calls to a particular tool than usual.
- Actions outside normal hours or on unusual records.
This is effectively anomaly detection applied to the AI itself.
4. Cost and load
Number of calls, volume of text processed, processing time. Not only for the budget: an agent caught in a loop often shows up first as a spike in cost.
Worked example: the price of a week without monitoring
Worked example: suppose an agent sends payment reminders for outstanding invoices. After a change in the accounting system, credit notes stop being included in the receivables list the agent reads, for a whole week. During that week the agent sends 60 reminders, 15 of them for amounts that have already been credited.
Without monitoring: the problem comes to light when customers call. Someone has to work out which of the 60 reminders were wrong, apologise to 15 customers and find the cause. Say this takes three days of work, plus damage to the relationship with a few good customers.
With monitoring: a daily check compares the total of the reminders with the outstanding balance in the ledger. After day one there is a discrepancy. The agent is paused, the cause is found, and 2 reminders are corrected instead of 15.
The difference is not in the quality of the agent, but in how quickly you see that its input changed. Source data errors like this are worked out in what happens when AI uses the wrong business data.
What makes a good log?
A log is only worth something if it meets a few conditions:
- Complete. Not just the outcome, but every step, including what the AI read.
- Unchangeable by the AI. A log the agent can edit or delete itself is not evidence. It belongs somewhere the agent can only write to.
- Kept long enough. Many problems only surface after weeks, at month-end close or when a customer calls. A seven-day log will be gone by then.
- Searchable. You must be able to find quickly everything the agent did for customer X, or everything it did between Tuesday and Thursday.
- Privacy-aware. Logs often contain the data the AI processed. If that includes personal data, the same rules apply as for the source. See how to protect personal data in AI systems.
Alerting, not just recording
A log nobody reads only helps after the fact. Monitoring becomes genuinely useful with alerts that go to a person automatically. Example rules:
- More than X write actions per hour: pause and alert.
- An action on a customer that was not part of the instruction: alert.
- A finding above a set euro amount: always to a person.
- The share of invalid findings rises above a threshold: alert the owner.
- A tool call that fails several times in a row: stop.
The design of those thresholds is tied to when a person should step in. That is covered in human-in-the-loop AI explained.
How do you set up AI monitoring? Step by step
- List every AI application that works with company data, including standalone tools that staff use.
- Decide for each application what it may do: read only, propose, or execute itself. The more it may do, the stricter the monitoring.
- Make sure every step is logged somewhere outside the application's own reach.
- Choose three quality measures to review monthly, such as the share of valid findings, deviation from the ledger and the number of manual corrections.
- Set up alert rules for unusual behaviour, with a named person who receives them.
- Test the emergency brake. Know who can pause an agent or revoke a key, and try it once.
- Evaluate after every model change whether the quality measures stayed the same.
Monitoring in a revenue intelligence system
In revenue control, monitoring has two sides. One is monitoring of the AI: is what the system reports correct? The other is monitoring by the AI: the system continuously watches for new leaks. RiOS has an AI monitoring module for the latter, which watches connected metrics continuously and raises a signal as soon as a figure deviates or a new leak appears. How both forms fit into the AI architecture for revenue control is explained in how AI works within Revenue Intelligence.
Frequently asked questions
Is logging the same as monitoring?
No. Logging records what happened. Monitoring actively looks at that data and alerts when something deviates. You need the first to be able to do the second.
How long should I keep AI logs?
Long enough to reconstruct a problem you only discover at month-end or quarter-end close. At the same time, no longer than necessary if they contain personal data. Set a retention period and be able to justify it.
Do I need monitoring if the AI only advises?
Yes, but lighter. You want to know whether the advice is right, otherwise people will either follow it blindly or ignore it. For AI that acts on its own, full step-by-step logging is essential.
Who in the business is responsible for AI monitoring?
The owner of the process the AI works in. For an agent in invoicing, that is the person responsible for invoicing, not the IT department alone.
More in this cluster
- How does AI work within Revenue Intelligence?Start here
- AI architecture for Revenue Intelligence
- What is anomaly detection?
- What is predictive analytics?
- Predictive AI vs generative AI
- What is an AI agent?
- What are AI agents in RevOps?
- What is tool calling?