How do you protect personal data in AI systems?
Protecting personal data in AI comes down to less data, tokenisation, agreements with processors and access control. A practical approach for B2B companies.
You protect personal data in AI systems by sending as little personal data as possible to the AI model, replacing the rest with tokens, enforcing per-user access in the data itself, recording the processing in a data processing agreement and knowing where the data ends up geographically. Data protection law such as the GDPR asks nothing new here: purpose limitation, data minimisation and appropriate security apply to AI just as they do to any other system. The difference is that AI makes it very tempting to send everything.
Why does AI increase the risk?
A B2B company often thinks of personal data as something for HR or consumer businesses. But a CRM and an accounting system are full of it: contacts, email addresses, phone numbers, notes on conversations, sometimes the home addresses of sole traders. With a sole trader or small partnership, even the company can be traced to a person.
AI changes three things:
- Volume. An analysis that used to run on a five-column export now receives the entire customer record, including free text, because the model "can read everything".
- Destination. The data leaves your systems and goes to a model at an external provider, sometimes outside the EU.
- Free text. Notes, emails and tickets suddenly become usable for analysis. That is exactly where data sits that nobody ever consciously recorded: a sick note, a personal situation, a dispute.
What does data protection law require when you use AI?
The GDPR contains a number of principles that translate directly to AI. Comparable laws elsewhere follow similar lines; check the rules in your jurisdiction.
Purpose limitation. You process personal data for a defined purpose. "AI analysis" is not a purpose. "Checking whether invoiced amounts match contractual arrangements" is. The purpose determines which data is needed.
Data minimisation. No more data than the purpose requires. Revenue control needs amounts, dates, contract terms and customer numbers. It rarely needs the names of contacts.
Integrity and confidentiality. Appropriate technical and organisational measures. The GDPR explicitly names pseudonymisation as an example of such a measure.
Processors. A party that processes personal data on your behalf, such as an AI provider or a software vendor using AI, is a processor. That requires a data processing agreement, which among other things lists the sub-processors involved.
Transfers outside the EU. If data goes to a country outside the European Economic Area, additional requirements apply. This affects many AI models, because the large providers are often based outside the EU or also process data there.
Automated decision-making. Decisions taken solely by automated means that significantly affect someone are subject to strict rules. In B2B revenue control this matters less, because it is usually about companies and a person takes the decision, but it is one more reason to let AI recommend rather than decide.
This is not legal advice. For your own situation it is sensible to consult your data protection officer or a lawyer, particularly if you are unsure whether a processing activity is high-risk. In that case a data protection impact assessment (DPIA) may be required.
Which measures actually protect personal data?
1. Less data to the model
The most effective protection is data the model never receives. For each AI application, make a list of the fields that are needed. Anything outside it is not sent. Free-text fields only if the task demonstrably needs them.
2. Tokenisation for what is sent
Identifying data that is needed for context, such as which invoices belong to the same contact, is replaced with tokens. The model sees PERSON_0412, not the name. The lookup table stays in your own environment. Worked out in how does tokenisation of business data work.
3. Access in the data, not only on the screen
An account manager who can only see their own customers in the CRM should also only be able to query their own customers through an AI assistant. That only works if the access rules live in the database itself, row-level security, and not only in the user interface. Otherwise a clever question to the AI can get round the interface.
4. Separate accounts with narrow permissions
An AI application gets its own account in each system, with read-only access to the data it needs. See how do you give AI access to business data.
5. Agreements with every processor
Record where the data is processed, which sub-processors are involved, whether input is stored or used to train models, and for how long. Consumer versions of AI tools often have different terms from business versions.
6. Logging and retention periods
Record which data was viewed by which application. And pay attention to the log files themselves: if an AI system's prompts and answers are logged, the personal data is in there too. A log file is also a form of processing, with a retention period.
Where does it go wrong in practice?
The pasted export. An employee pastes a customer export into a free chat tool for a quick analysis. Names, email addresses and revenue per customer now sit with a provider with whom no data processing agreement has been signed. It is the most common and most underestimated route.
Notes nobody read. An AI application that summarises CRM notes suddenly surfaces information that sat unseen in free-text fields for years. Sometimes sensitive, sometimes wrong, sometimes both.
Prompts in log files. The AI system itself is well protected, but the full prompts, including customer data, are logged in a monitoring tool that half the IT department can access.
Sub-processors that change. A vendor switches AI model or adds one. The data now goes to another party, perhaps in another country, without anyone noticing.
Data that stays behind. After a trial or a contract ends, customer data is still with the vendor, because nobody asked for it to be deleted.
Worked example: how much personal data does a revenue analysis need?
Worked example: suppose you want to check whether all additional work on projects has been invoiced. For each project you have a CRM record with 40 fields, 12 of which contain personal data: the customer's project manager, billing contact, phone numbers, email addresses and a notes field.
The analysis needs: project number, customer number, original order value, registered change orders with amount and date, and invoiced amounts. That is 6 fields, none of which contains personal data.
- Without selection, 40 fields per project are sent, 12 of them with personal data.
- With selection, 6 fields are sent, without personal data.
- Only if you want to find additional work that is recorded solely in notes do you need the notes field. Then you tokenise the names in it.
The result of the analysis is the same. The risk is a fraction. How to find the leak itself is covered in how do you find forgotten invoices.
Checklist for every AI application
- Is the purpose described in one sentence?
- Which fields are needed for that purpose, and which of them contain personal data?
- Is identifying data tokenised before it goes to the model?
- Where does the model run, and where is data stored?
- Is there a data processing agreement, with a list of sub-processors?
- Is input used for training? If so, can that be switched off?
- Are prompts and answers logged, and who can view those logs?
- What happens to the data when the contract ends?
- Is per-user access enforced in the data itself?
- Has it been assessed whether a DPIA is needed?
How does RiOS handle this?
RiOS is being built with a number of fixed design rules that address this directly: a tokenisation layer that removes personal and identifying data before anything reaches an AI model, row-level security in the database, processing on EU-based infrastructure, and a data processing agreement naming every sub-processor before anything is connected. The details are on the security page. How privacy fits into the rest of the AI architecture for revenue control is covered in how does AI work within Revenue Intelligence.
Frequently asked questions
Is using AI on customer data allowed under the GDPR?
Yes, if you meet the usual requirements: a purpose and a legal basis, no more data than necessary, appropriate security and agreements with processors. AI is not an exception, but it does make it easier to process too much by accident.
Are business contact details also personal data?
Yes. A name with a business email address can be traced to a person and falls under the GDPR. The details of a limited company are not personal data, but those of a sole trader can be.
Is an AI model within the EU enough?
It solves the question of international transfers, not that of data minimisation. Even a model within the EU usually does not need to see your customers' names.
What is the quickest improvement I can make tomorrow?
Agree that no customer exports are pasted into chat tools without a data processing agreement, and make a business environment available for anyone who wants to use AI.
More in this cluster
- How does AI work within Revenue Intelligence?Start here
- AI architecture for Revenue Intelligence
- What is anomaly detection?
- What is predictive analytics?
- Predictive AI vs generative AI
- What is an AI agent?
- What are AI agents in RevOps?
- What is tool calling?