AutoMaat
Knowledge base· AI technology

How does tokenisation of business data work?

Tokenisation replaces names, email addresses and other identifying data with tokens before AI sees them. How it works and what it does and does not protect.

Ricardo Mastenbroek7 min read
Lees dit artikel in het Nederlands

Tokenisation of business data means that identifying information, such as names, email addresses, phone numbers and bank account numbers, is replaced by neutral placeholders, tokens, before the data is sent to an AI model. The model works with the tokens, and only when the result comes back into your own environment are they translated back into the real data. This lets AI analyse patterns in your revenue data without ever seeing who it concerns.

What does "token" mean here?

First, a possible confusion. In AI, "token" also means something else: the piece of text a language model splits its input into, roughly part of a word. When a provider quotes prices per thousand tokens, that is the meaning they have in mind.

This article is about the privacy meaning: a placeholder for a sensitive piece of data. "John Smith, [email protected]" becomes, for example, "PERSON_0412, EMAIL_0412". The concept comes from the payments world, where credit card numbers have long been replaced in this way.

How does tokenisation work?

The process has four steps.

1. Recognise. A processing step runs through the data and recognises what is identifying. In structured fields that is straightforward: the "contact" column contains names, the "email" column email addresses. In free text, such as a CRM note or an email, it is harder. That requires pattern recognition and language recognition.

2. Replace. Every recognised item is replaced by a token. Importantly, the same item always gets the same token. If John Smith appears on ten invoices, he is PERSON_0412 on all ten. That keeps patterns visible: the model sees the same person recurring ten times, without knowing who it is.

3. Process. The AI model receives only the tokenised version. It analyses, compares, summarises and writes an answer containing tokens.

4. Translate back. The lookup table, which token belongs to which real item, stays in your own secure environment. The model's answer is converted back into readable data there, for the user who is allowed to see it.

An example of what is sent to the model:

Original Sent to the model
Customer: Example Engineering Ltd Customer: CUSTOMER_0087
Contact: Sarah White, [email protected] Contact: PERSON_1203, EMAIL_1203
Note: "Sarah agreed to 5% discount until Dec" Note: "PERSON_1203 agreed to 5% discount until Dec"
Contract value: EUR 64,000 Contract value: EUR 64,000

Amounts, dates and arrangements stay as they are. The model needs those to find leakage. It does not need to know who the people are.

What should you tokenise and what not?

The trade-off is always: does the model need this item to carry out the task?

Almost always tokenise:

  • Names of people
  • Email addresses and phone numbers
  • Addresses of individuals
  • Bank account numbers
  • National identification numbers and similar identifiers (which as a rule should not be in your CRM or accounts in the first place)

Depending on the task:

  • Company names. A company name is usually not personal data, but for a sole trader or a small partnership it can point directly to a person. It can also be commercially sensitive.
  • Product names and project names, if they can be traced to a customer.

Usually do not tokenise:

  • Amounts, quantities, prices, percentages
  • Dates and terms
  • Product categories and contract types
  • Status fields

This last group is exactly what revenue control revolves around. Tokenise that as well, and the model can no longer do anything.

What is the difference between tokenisation, pseudonymisation and anonymisation?

Three terms that are often used interchangeably, with different legal consequences.

Anonymisation means that data can irreversibly no longer be traced to a person, not even in combination with other information. Truly anonymised data falls outside data protection law such as the GDPR. In practice genuine anonymisation is difficult, especially with small datasets.

Pseudonymisation means that data can only be traced back with additional information, such as a lookup table, and that this information is stored separately and securely. Pseudonymised data remains subject to the GDPR, but the risk is lower, and the GDPR explicitly names pseudonymisation as an appropriate security measure.

Tokenisation with a lookup table is a technique for pseudonymisation. For the AI model itself, which never sees the lookup table, the data is not traceable. For your organisation, which manages the table, it is.

The practical consequence: tokenisation reduces the risk considerably, but does not release you from your data protection obligations. Check the rules in your jurisdiction. This is worked out further in the article on how do you protect personal data in AI systems.

Where does tokenisation go wrong?

Free text. Structured fields can be tokenised reliably. A CRM note such as "spoke to the owner's wife, she has just had an operation, call after the summer" contains sensitive information without a single name. No recogniser catches everything. That is why it is sensible to send free text only when the task genuinely needs it.

Identifiable through combinations. A token hides the name, but a combination of data can still make someone recognisable. "The only customer in one particular region with a contract above EUR 500,000" is identifiable without a name. With small customer bases this weighs more heavily.

Inconsistent tokens. If "J. Smith", "John Smith" and "[email protected]" get three different tokens, the model sees three people. Customer-level analyses then become unreliable. Good tokenisation normalises first.

The lookup table in the wrong place. If the table sits in the same environment as the model, or ends up in a log file, the protection is gone.

Tokenisation as an excuse. "It is tokenised, so we can send everything" is the wrong reflex. Data minimisation remains the starting point: send only what is needed, and tokenise whatever of that is identifying.

Worked example: does leak detection still work after tokenisation?

Worked example: suppose you want to check whether customers with a discount arrangement go back to paying the full rate once that arrangement ends. You have 1,200 invoice lines and 85 CRM notes with discount arrangements.

After tokenisation the model sees a token for each customer, a percentage and an end date for each arrangement, and an amount and a date for each invoice. It finds, for example, that CUSTOMER_0087 had a 5 percent discount until the end of December, and was still invoiced at the discounted rate from January to March, on a monthly amount of EUR 5,300.

  • Missed revenue: 5 percent of EUR 5,300 is EUR 265 a month, EUR 795 over three months.
  • For the employee, CUSTOMER_0087 is translated back to Example Engineering Ltd, so the invoice can be corrected.

The model found the discrepancy without knowing which customer it concerned. That is the heart of it: leak detection needs amounts, dates and arrangements, not names. More on this type of leak in revenue leakage from wrong prices.

Which questions should you ask your vendor?

  1. Is data tokenised before it goes to an AI model? For every model, or only some?
  2. Which data is recognised? Only structured fields, or free text as well?
  3. Where is the lookup table kept, and who can access it?
  4. Does the same item always get the same token?
  5. What is sent without tokenisation, and why?
  6. Has the tokenisation been tested, for example by sampling what the model actually receives?

How does RiOS use tokenisation?

RiOS is being built with a tokenisation layer as a fixed design rule: a processing step removes names, email addresses and other identifying data and replaces them with tokens before anything reaches an AI model. That applies to every model, wherever it runs. The data flows are explained on the security page. How tokenisation relates to access permissions and connections is covered in how do you give AI access to business data, and the wider picture in how does AI work within Revenue Intelligence.

Frequently asked questions

Does tokenisation make data anonymous?

Not in the legal sense. As long as a lookup table exists, it is pseudonymisation. For the AI model the data is not traceable, but data protection law still applies to your own processing.

Does it make the AI less capable?

For revenue analysis hardly at all, because amounts, dates and arrangements remain. For tasks where the name itself matters, such as writing a personalised email, the name is inserted again only after the model has done its work.

Is tokenisation the same as encryption?

No. Encrypted data is unreadable until you decrypt it, and a model can do nothing with it. Tokenised data is readable and usable; only the identifying parts have been replaced.

Do I need to tokenise if the AI model runs in the EU?

Where the model runs does not change the fact that it does not need your customers' names for most analyses. Data minimisation applies everywhere.

Share this article
Knowledge base · AI technology

More in this cluster

All 22 topics in this cluster

More from AutoMaat

Rather know what this costs you specifically?

The Revenue Audit puts a euro amount on where your revenue leaks.

Plan the Revenue Audit