Revenue data architecture for B2B
How to structure a B2B company's revenue data: source systems, keys, a shared revenue model and checks, without a large data project.
A revenue data architecture is the way your company's revenue data is spread across systems and connected: which systems record which step, which keys link them, which system leads for each piece of data, and where you bring them together to compare. For a B2B company with EUR 2 million or more in revenue, the goal is not one big system but a chain of sources that can be traced from lead to paid invoice. A good architecture shows where revenue disappears. A bad one mainly produces attractive dashboards.
Why should a director care about this?
"Architecture" sounds like something for the IT department. But the questions it answers are management questions. What revenue should we expect? What have we invoiced? Why do they differ? Which customers are buying less than agreed? If you cannot get an answer to those without someone spending three days in Excel, your architecture is the problem, not your people.
Most B2B companies have no architecture that was ever designed. They have a collection of systems that grew over the years: an accounting package from the early days, a CRM that sales once chose, a planning tool from operations, a time-tracking system, and Excel in between. That is normal. It becomes a problem when nobody knows how it all fits together.
What are the layers of a revenue data architecture?
A usable revenue data architecture for B2B consists of five layers. You probably have all of them already, just not explicitly.
Layer 1: source systems
The systems where a step in the revenue chain is recorded. For a typical B2B company:
| Step | Typical system | Examples |
|---|---|---|
| Lead and deal | CRM | Salesforce, HubSpot, Pipedrive |
| Quote and price | CRM, CPQ or ERP | Quoting module, price lists in the ERP |
| Contract | Contract module, document folder | Signed PDFs, contract management |
| Order and delivery | ERP, planning | SAP, NetSuite, Dynamics, Exact, AFAS |
| Hours and projects | Time tracking, project module | Often part of the ERP |
| Invoice and payment | Accounting | Xero, NetSuite, Exact, Twinfield |
| Usage and service | Product data, support system | Licence management, ticketing |
The first job is to fill in this table for your own company. Including the Excel files that are in reality a source system. A price agreement list in Excel used by sales support is a source, whether you like it or not.
Layer 2: keys
The fields that connect a record in one system with a record in another. Customer number, deal number, contract number, order number, item number. Without keys, your source systems are islands. In almost every company this is the weakest layer, and the cheapest to improve. The article how to connect sales data with financial data explains how to put them in place.
Layer 3: ownership per piece of data
One leading system for each piece of data. Contract value: the contract. Delivered: the ERP. Invoiced: the accounting system. Expected: the CRM. Without this agreement, every difference becomes a debate about who is right instead of a question about what went wrong. See which data source leads for revenue.
Layer 4: a shared revenue model
A place where the sources come together in one model with shared definitions. What is a customer? What is recurring revenue? When is a deal won? How do you count a multi-year contract? This model can live in a data warehouse, in a BI environment or in a platform that reads the sources. What matters most is not where it lives, but that the definitions are written down and everyone uses the same ones.
Layer 5: checks and signals
Rules that run over the revenue model and report exceptions. Won deal with no order. Order with no invoice. Contract with no indexation. Customer with declining purchases. This is the layer that delivers money. The first four layers exist to make it possible.
Many companies stop at layer 4: they build a data warehouse and dashboards, and then someone looks at charts every month. A dashboard shows a total. A check shows an exception, with an owner. For revenue leakage you need the second.
What are the architecture choices?
There are three common ways to set up layers 4 and 5.
Everything in the ERP. Some companies choose to do as much as possible in one ERP, with CRM, projects and invoicing as modules. That reduces the number of integrations. It does not solve the fact that an agreement in the CRM module does not automatically end up in the invoicing module. Even within one package you need keys and checks.
A data warehouse with BI. Sources are copied daily to a central database, for example Snowflake, and you build reports on top in Power BI or a similar tool. This gives a lot of freedom. It also needs people to build and maintain it, and it often gets stuck at reporting rather than checking. The difference is set out in Revenue Intelligence vs data warehouse.
A platform that reads the connected sources. A Revenue Intelligence platform connects to the existing systems, builds the shared model itself and runs the checks. Your source systems remain leading and you have less to build yourself. You do depend on which systems the platform can read.
These choices are not mutually exclusive. A company with a data warehouse can put a checking layer on top of it. What does not work is making no choice: then Excel becomes your architecture.
Principles that apply to every choice
- Source systems remain leading. The revenue model is a copy and a comparison, not a new source. Corrections are made in the source, not in the model.
- Read before you write. Start with read-only access to the sources. Writing back, such as changing a record or creating a task, only once the comparison is reliable and someone approves it.
- Keep history. A CRM overwrites fields. If a contract value changes, you want to know what it was. Keep snapshots, so you can see when an exception arose.
- Settle definitions before you build. A dashboard without agreed definitions produces a new debate, not an answer.
- Start small, at the most expensive seam. Not all sources at once. Start where the most money sits between two systems.
Where does it go wrong?
Starting with the technology. A project that begins with the question of which data warehouse to choose often ends with a data warehouse and few answers. Start with the questions you want answered and the keys you need.
Not planning maintenance. Source systems change. Fields are renamed, products added, packages replaced. An architecture without an owner goes quiet within a year, and nobody notices until a report shows a suspiciously good number.
Forgetting personal data. A revenue model contains contacts, email addresses and sometimes more. Decide in advance which data you really need, where it is stored and who has access. For most checks you do not need the names of contacts, only customer and contract data.
Worked example
Worked example: suppose a business services company with EUR 15 million in revenue is considering two routes. Route A is its own data warehouse with dashboards, built and maintained by an external partner. Route B is first putting keys in place and running five checks on the most expensive seams. Which route is better does not depend on the cost of the technology, but on what it delivers. If the checks in route B find EUR 90,000 in missed invoicing in the first year, and route A mainly delivers nicer monthly reports in that same year, the choice is clear. If the data warehouse already exists, the cheapest step is often the checking layer on top.
The amounts in this example are illustrative. The lesson is that you judge an architecture by what it finds, not by what it shows.
Checklist: how does your architecture stand?
- Is there an overview of all systems, including Excel files, that record a step in the revenue chain?
- Does every customer in the CRM carry the debtor number from the accounting system?
- Can a won deal be traced to the order or contract that came out of it?
- Has it been agreed for each piece of data which system leads?
- Are the definitions of revenue, customer and recurring revenue written down?
- Is at least one automatic check running on a seam between two systems?
- Is there an owner who maintains the integrations and checks?
Fewer than four yeses means the architecture exists mainly out of habit. The pillar how do you get a single source of truth for revenue? describes how to build on from there.
Frequently asked questions
Do I need a data warehouse?
Not necessarily. A data warehouse is one way to set up layer 4. For many companies with EUR 2 to 20 million in revenue it is a heavy investment for what you actually need: keys, definitions and a few checks.
Who owns the revenue data architecture?
Ideally the person responsible for finance, with a clear role for sales operations or IT for maintenance. What matters most is that it is one person who is allowed to decide on definitions across departments.
Do we have to clean our data first?
Only the data you need for your first checks. A full clean-up project in advance takes a long time and delivers nothing visible. Checks show which data really matters.
Where does AI fit in this architecture?
AI is useful in layer 2 and layer 5: matching records that have no key, reading terms from contracts and spotting unusual behaviour. It does not replace keys or definitions. An AI analysis on an architecture without agreements gives quick answers to the wrong questions.
More in this cluster
- How do you get a single source of truth for revenue?Start here
- CRM vs ERP: where does your real revenue come from?
- Why CRM data is not the same as financial data
- CRM-to-billing reconciliation explained
- How do you connect CRM to billing?
- How do you connect CRM to ERP?
- How do you check CRM data automatically?
- How do you check billing automatically?