AutoMaat
Knowledge base· AI technology

What is anomaly detection?

What anomaly detection is, how it finds deviations in revenue data, why it misses some revenue leaks and how to keep it from producing only noise.

Ricardo Mastenbroek8 min read
Lees dit artikel in het Nederlands

Anomaly detection is the automatic identification of data that deviates from what is normal. A system first learns the usual pattern, for example how much a customer buys each month or what discount is typically given on a product, and then flags whatever falls clearly outside it. In revenue work this is how you find customers who are quietly buying less, discounts that have suddenly crept up and invoices that do not match the contract. It only finds what changes: a leak that has always been there looks normal to it.

How does anomaly detection work?

The principle has three steps.

  1. Establish the normal pattern. The system looks at the history: what is the usual value, and how much does it normally fluctuate? For a customer who has ordered every month for two years, that is an average monthly spend with a certain spread.
  2. Compare new values. Every new month, invoice or order is set against that pattern.
  3. Flag what deviates too far. When a value lies further from normal than a threshold, it becomes a signal.

The hard part is the first step. "Normal" is not a single number. A customer who always buys little in December is not unusual in December. A new customer has no pattern yet. A growing business has a normal that shifts every year. How well anomaly detection works depends on how intelligently normal is defined.

What types of anomaly are there?

Point outliers. A single value that deviates sharply: an invoice of EUR 48,000 for a customer who normally sits around EUR 4,800. Often a typing error, sometimes a genuine one-off order.

Contextual anomalies. A value that is normal in itself, but not at that moment or for that customer. Revenue of EUR 80,000 is ordinary for November, but low for March if March is always your strongest month. A 15 percent discount is normal for key accounts, but not for a customer who spends EUR 5,000 a year.

Collective anomalies. No single value is extreme, but together they form a pattern. Three months in a row of slightly lower purchases. A payment term that stretches by a few days every month. For revenue this is often the most important type, because customers rarely leave all at once.

Which methods are used, in plain language?

Fixed thresholds

The simplest form: flag everything above or below a limit. Discount above 20 percent, invoice above EUR 50,000, payment term above 60 days. Clear and explainable, but one limit for every customer and product produces many false alarms on large customers and misses deviations on small ones.

Statistical deviation per series

For each customer, product or metric you determine the average and the normal spread. A value that deviates by more than a set multiple of that spread is a signal. This is the core of most anomaly detection on revenue data.

Worked example: suppose a customer has bought an average of EUR 12,000 a month over the past 24 months, with a normal fluctuation (standard deviation) of EUR 1,000. The last three months came in at EUR 9,500, EUR 9,000 and EUR 8,800. The middle month is EUR 3,000 below the average, three times the normal fluctuation. One such month can be chance: a holiday, a postponed order. Three months in a row well below the normal range almost never is. If it stays that way, that is roughly EUR 36,000 of revenue a year disappearing, and often the start of a customer who is leaving. How to spot these customers early is covered in how do you spot customers who are quietly buying less.

Removing season and trend

For series with a seasonal pattern you do not compare with the overall average, but with what is normal for that month, adjusted for growth. You compare December with previous Decembers, not with October. Without that adjustment you get a wave of false alerts in the same months every year.

Comparing with peers

Not everything has its own history. A new customer, a new product. In that case you compare with similar customers or products. A customer on a 30 percent discount while comparable customers sit around 10 percent is an anomaly, even if it has been that way from day one. This is the method that finds leaks that were always there.

Machine learning

Models such as isolation forests or autoencoders look for anomalies across many attributes at once: a combination of amount, discount, product, customer type and salesperson that rarely occurs together. They find patterns a person would not have thought of, but they are harder to explain. On revenue data they work best as a complement to the simpler methods, not as a replacement.

What does anomaly detection find in B2B revenue data?

  • Spend per customer falls gradually over a few months.
  • Discount per product or salesperson is clearly higher than normal in a quarter.
  • Invoice amount deviates from the contract: the price has not been indexed, or a licence extension is missing from the invoice.
  • Number of credit notes for a customer or product group rises.
  • Payment term of a customer keeps getting longer, often an early sign of problems or dissatisfaction.
  • Additional work per project is strikingly low compared with similar projects, which may mean it was carried out but not invoiced.
  • Product mix shifts: a customer still buys the same volume, but only the cheaper products.

Why does anomaly detection not find every leak?

This is the most important point for anyone using it against revenue leakage. Anomaly detection finds change. A leak that was there from the start is normal as far as the system is concerned.

A customer who has been invoiced without indexation for three years has a perfectly stable pattern. A contract with one component that never made it onto the invoice produces the same amount every month. No deviation, so no alert.

That is why you need two other methods alongside anomaly detection: rules that compare what should be there with what is there (contract against invoice, won deal against order), and comparison with peers. Together they find far more than any one of them alone. How these methods work together is explained in how do you detect revenue leakage automatically.

And the reverse: not every anomaly is a leak. A customer buying less may simply need less. A high discount may have been given deliberately for a multi-year deal. An anomaly is a question, not a conclusion.

Where does it go wrong in practice?

Too many alerts. The biggest risk. A system that flags two hundred anomalies every week is ignored after a month. The result is worse than having no system, because the real signals are lost in the noise.

Thresholds that do not fit. A threshold that is right for a customer worth EUR 500,000 a year is far too coarse for a customer worth EUR 5,000. Thresholds should be set per series, not for everything at once.

Seasonality ignored. The same wave of alerts every year in summer and around the holidays.

Normal shifts. When a business grows, raises a price or launches a new product, what counts as normal changes. A model that is not updated keeps measuring against old patterns.

No euros. An anomaly of three standard deviations means nothing to a finance director. "This customer has been buying EUR 3,000 a month less since July" does.

How do you make anomaly detection useful?

  1. Rank by euro impact. Do not show every anomaly, but the anomalies ranked by what they cost if they persist. The top ten is often enough.
  2. Combine signals. A customer with falling spend, a longer payment term and more support tickets is a stronger signal than any of these alone.
  3. Show the source. Which invoices, which months, which comparison. A person must be able to see within a minute whether it holds up.
  4. Show the confidence. A signal based on 24 months of history is more certain than one based on four. How to judge that confidence is covered in how do you know whether an AI recommendation is reliable.
  5. Learn from the response. When an alert is closed as "explained", similar noise should come back less often.
  6. Give every alert an owner. An anomaly without someone to investigate it remains just a number.

Where does it fit?

Anomaly detection is one layer in a Revenue Intelligence system, alongside rules, predictive models and language models. How those layers work together is explained in how does AI work within Revenue Intelligence. Whether AI in general can find revenue leakage, and where its limits lie, is covered in can AI detect revenue leakage.

Frequently asked questions

How much history do you need?

For a series without seasonality, twelve data points is a reasonable minimum, for example twelve months. For a seasonal pattern you need at least two years. For new customers or products you use comparison with peers.

Can you do anomaly detection in Excel?

For a limited number of customers, yes: calculate the average and standard deviation per customer and highlight anything that deviates by more than two or three times the spread. It becomes unworkable once you have many series, seasons and combinations of signals.

What is the difference between anomaly detection and a rule?

A rule checks something you know in advance: every won deal must have an invoice. Anomaly detection flags something that deviates from normal, without you knowing beforehand what you were looking for. You need both.

How do you stop the system from flagging too much?

By setting thresholds per series, taking seasonality into account, combining signals and only showing the anomalies with the largest euro impact. Fewer, better alerts are worth more than completeness.

Share this article
Knowledge base · AI technology

More in this cluster

All 22 topics in this cluster

More from AutoMaat

Rather know what this costs you specifically?

The Revenue Audit puts a euro amount on where your revenue leaks.

Plan the Revenue Audit