AI architecture for Revenue Intelligence
How a Revenue Intelligence system is built technically, from connectors and data model to agents and audit log, and which choices make it reliable.
The architecture of a Revenue Intelligence system consists of layers that each have one job: connectors that read data from CRM, billing and ERP, a common data model that links customers, contracts and invoices, a layer of fixed definitions and rules, models that find anomalies and make predictions, language models and agents that explain and act, and around all of it privacy, access rights, approval and an audit log. The AI models are the smallest part. Whether the system is reliable is determined by the data model, the access rights and the controls around it.
Why architecture matters more here than the model
Revenue Intelligence is about money. A finding such as "this customer pays EUR 3,800 a year too little" has to be right, otherwise you send a customer an unjustified correction or miss a real leak. That calls for a system in which you can see, for every figure, where it came from, which steps it went through and who approved it.
A standalone chatbot on your CRM cannot do that. A well-built architecture can. The overview article how does AI work within Revenue Intelligence describes the layers in terms of what they do. This one is about how they fit together technically and which choices you make along the way.
What are the layers of a Revenue Intelligence architecture?
| Layer | Job | Key design choice |
|---|---|---|
| Connectors | Read data from source systems | Read-only, revocable per system |
| Raw storage | Keep data as it arrived | Keep history, do not overwrite |
| Common data model | Link customer, contract, order, invoice, deal | One key per customer across all systems |
| Definitions and metrics | Fixed calculation of revenue, margin, retention | One definition, recorded, not per report |
| Detection and prediction | Rules, anomaly detection, predictive models | Every result with source and confidence |
| Privacy layer | Replace personal data before a language model | No exceptions |
| Language models and agents | Read, explain, prepare actions | Access through fixed tools, not the raw database |
| Approval and execution | Write actions back to source systems | Only after approval by a person |
| Access and audit log | Who sees what, what happened | Rights in the database, log out of agents' reach |
| Monitoring | Watch that everything still holds | Alerts on connectors, data quality and models |
Connectors: read, do not take over
The first layer fetches data from the systems you already have: Salesforce, HubSpot or Pipedrive for deals; Xero, NetSuite, SAP, Dynamics, Exact or AFAS for invoices and debtors; a document folder for contracts; Zendesk or another support system for tickets. Which systems you need at minimum is covered in which systems need to be connected for Revenue Intelligence.
Three choices determine the quality of this layer.
Read rights. To find revenue leaks, the system does not need to change anything. A connector with read-only rights cannot break anything, even if another part of the system makes a mistake. Write rights belong separately, in the execution layer, and only for what is needed.
Incremental fetching. Re-fetching ten years of invoices every night is slow and puts load on your systems. A good connector fetches only what has changed since last time, and knows what to do when something has been deleted.
Keeping history. A CRM overwrites the close date when a salesperson changes it. For a forecast and for finding patterns, you want to know precisely that the date has been moved three times. Raw storage therefore keeps every version, not just the latest.
The common data model: where most of the work is
This is the least visible layer and the one that determines the most. Revenue leakage sits between systems, so the system has to know which records in different systems refer to the same thing.
A common data model defines the core objects, usually customer, contract, product, order, invoice and deal, and records for each object what it is called in every source system. Customer K-1042 is account 00158 in the CRM, debtor 20417 in the accounting system and the folder "Jansen Construction" in the contract store. That translation is called identity resolution, and it is where most errors occur.
An example of what happens then: if two branches of the same customer are treated as two customers, one branch looks like a customer shrinking sharply and the other like a new customer. The system reports churn that does not exist. Conversely, if two different customers with similar names are merged, real anomalies disappear into the average.
Good identity resolution uses hard keys where they exist (company registration number, VAT number, a customer number present in both systems) and proposes the rest for confirmation. A model can suggest that two records belong together; a person or a fixed rule confirms it. More on building such a data model in revenue data architecture for B2B.
Definitions: record once what revenue is
"Revenue" means something different in the CRM than in the accounts, and "active customer" means something different to sales than to finance. If every report and every model uses its own definition, you get dashboards that contradict each other.
A definitions layer, also called a semantic layer, records each metric once: what counts as realised revenue, when a customer is lost, how a credit note is offset, on which date an annual contract counts. All rules, models, dashboards and language models use that one definition. It also stops a language model from inventing its own calculation: it asks the definitions layer for the figure.
Detection and prediction
This is where the rules and models that find anomalies run. Rules for what you already know: a won deal without an invoice, a contract with an indexation clause without a price increase. Anomaly detection for what deviates from the normal pattern. Predictive models for probabilities of winning, churn and revenue.
The design choice that matters most: every result carries its source and its confidence. A finding is not just "customer K-1042 is paying too little", but also which invoice lines and which contract, which rule or which model, and how certain. Without that metadata a finding cannot be checked, and therefore cannot be trusted.
A second choice is what happens to findings with low confidence. A system can show them with a warning, or not publish them at all until there is more evidence. The second is called fail-closed: when in doubt, show nothing. It prevents a dashboard full of half-signals that nobody looks at any more.
The privacy layer in front of every language model
Language models often run with an external provider. Whatever you send to such a model leaves your own environment. For revenue analysis that is usually unnecessary: the model does not need to know your contacts' names to read a contract clause or explain a finding.
A privacy layer replaces names, email addresses, phone numbers and other identifying data with tokens before anything reaches a model, and restores them when the answer is displayed within your own environment. The model works with "contact T-81 of customer K-1042"; your user sees the real name. How that works and where it gets difficult is covered in how does tokenisation of business data work.
Language models and agents: tools, not a bunch of keys
In a good architecture, a language model or agent gets no direct access to the database. It gets a fixed set of tools: "find the invoices for this customer", "return the clauses of this contract", "calculate revenue according to the standard definition". Each tool does one thing, with the rights that go with it.
That has two advantages. The model cannot do anything that does not exist as a tool. And every figure the model mentions comes from an exact query on the definitions layer, not from a guess. For questions about figures that is always the right route; searching documents with a language model is for text. That distinction is explained in RAG vs database queries for business data.
Approval and execution
A system that finds leaks but does nothing with them becomes a report that ends up in a drawer. A system that changes everything itself becomes a risk. The architecture in between: actions are prepared as proposals, with the finding and the source attached, and only carried out after approval by someone with the authority to give it. On approval, a separate execution layer, with only the write rights that action needs, writes back to the source system.
For some actions the approval can be relaxed over time, for example creating a task in the CRM. For others it always remains necessary, for example anything that reaches a customer.
Access, audit log and monitoring
Access. Who may see which customers, amounts and findings is most safely enforced in the database itself, with row-level security, and not only in the application. Then a bug in the front end cannot bypass it.
Audit log. For every step: which data was read, which model or rule produced a finding, who approved it, what was written back. The log sits somewhere agents can only append to, not modify.
Monitoring. A connector missing a field since an update, a model whose accuracy is declining, an agent suddenly proposing many more actions. Without monitoring you only notice when the figures stop adding up. See AI monitoring explained.
Choices you make in advance
Batch or real time? For most revenue leaks, nightly processing is enough. An invoice that appears as an anomaly a day late costs nothing. Real time is more expensive and more complex, and only needed for things like fraud or stock.
Your own data warehouse or a separate platform? If you already have a data warehouse, for example in Snowflake, you can build the lower layers on it. You then have to build and maintain the data model, definitions, rules and controls yourself. A separate platform brings those with it, but you depend on how it is set up.
Which models where? Predictive models often run perfectly well in your own environment. Large language models usually do not, and that is what the privacy layer is for.
For RiOS, the target architecture, with a read layer, tokenisation layer, row-level security and fail-closed publication, is set out on the security page. The platform is in beta; the page describes the model it is being built to.
Checklist for an architecture conversation
- Which connectors are read-only, and which have write rights, under which account?
- Is the history of changes kept, or only the latest state?
- How are customers matched across systems, and who confirms doubtful cases?
- Where are the definitions of revenue and retention recorded, and do all components use the same ones?
- Does every finding have a source and a confidence, and what happens below the threshold?
- What does a language model see of your customers?
- Can an agent do anything outside its tools?
- What is logged, where, and who can change it?
- Who gets an alert when a connector or model starts performing worse?
Frequently asked questions
Isn't this just a data warehouse with AI on top?
The foundation looks similar. The difference lies in what sits on top: fixed definitions of revenue, rules and models for leaks, a privacy layer, tools for agents, approval and an audit log. A data warehouse without those layers is storage, not Revenue Intelligence.
Does all the data have to go to one place?
For analysis, yes, a copy of the relevant fields; otherwise you cannot put systems side by side. Your source systems remain the source of truth; the system reads them and does not replace them.
How long does it take to set up?
That depends mainly on the state of your data and the number of systems, not on the AI. Matching customers across systems and recording definitions takes the most time.
Where does it most often go wrong in practice?
In identity resolution and definitions. A model running on data in which customers are wrongly matched or revenue is calculated three different ways produces convincing but wrong findings.
More in this cluster
- How does AI work within Revenue Intelligence?Start here
- What is anomaly detection?
- What is predictive analytics?
- Predictive AI vs generative AI
- What is an AI agent?
- What are AI agents in RevOps?
- What is tool calling?
- How do you give AI access to business data?