AutoMaat
Knowledge base· AI technology

Embeddings explained for business software

What embeddings are, how AI uses them to find text by meaning, and why amounts and contract terms should never rest on them.

Ricardo Mastenbroek9 min read
Lees dit artikel in het Nederlands

An embedding is a series of numbers that captures the meaning of a piece of text, so that software can calculate which texts resemble each other. Two sentences that mean the same thing but use different words get numbers that sit close together. That lets business software search on meaning rather than on exact words. For amounts, quantities and dates, embeddings are precisely the wrong tool: those you retrieve from a database, not from a similarity score.

What is an embedding, exactly?

An embedding model is an AI model that does only one thing: it reads a piece of text and returns a long row of numbers. That row is called a vector. Depending on the model, it contains hundreds to a few thousand numbers. Each number on its own means nothing to a person. Together they form a position in a space with a very large number of dimensions.

The useful part is the distance. Texts with a similar meaning end up close together. "The customer wants to cancel the contract" and "Customer indicates they will not renew" share almost no words, but their vectors lie close together. "The customer wants to extend the contract" lies further away, even though the sentence looks more like the first one word for word.

That distance is usually calculated with a simple formula, often called cosine similarity. The result is a number that says how strongly two texts resemble each other. That is all it is. An embedding does not know what is true, does not know what an invoice is and cannot calculate. It places text on a map.

What are embeddings used for in business software?

In business software you mainly come across embeddings in four places.

1. Searching by meaning

An employee searches the support system for "customer cannot log in". An ordinary search only finds tickets containing those words. Searching with embeddings also finds "password no longer works" and "access blocked". This is called semantic search.

2. Finding documents for an AI assistant

When a language model has to answer a question about your contracts or manuals, the right passage has to be found first. To do that, documents are cut into chunks, an embedding is made of each chunk and they are stored in a vector database. When a question comes in, the system looks for the chunks closest to the question and passes them to the model. That is the core of RAG, retrieval-augmented generation. How that compares with querying your database directly is covered in RAG vs database queries for business data.

The same company appears in your CRM as "Baker Installation Services Ltd", "Baker Installations" and "Baker IT Ltd". An exact comparison sees three customers. Embeddings of name, address and description show that two of them are probably the same. Note the word probably: "Baker IT Ltd" may be a different company. A person or a hard rule (company registration number, VAT number) has to settle it.

4. Classifying free text

Account manager notes, cancellation reasons, descriptions of additional work on a job sheet. This is text nobody has put in a fixed field. With embeddings you can group it: which cancellations are about price, which about service, which about an acquisition. That gives you structure in what was never recorded in a structured way.

Where it goes wrong: embeddings and numbers

This is the trap for anyone working with revenue. Embeddings are made to capture meaning, not to preserve precision.

To an embedding model, "Indexation 3.2 percent from 1 January" and "Indexation 5.2 percent from 1 January" look almost identical. They are two sentences on the same subject with one different figure. For your revenue, that difference is exactly what matters. The same applies to "licence for 40 users" versus "licence for 400 users", or "invoice quarterly in advance" versus "invoice quarterly in arrears".

Three consequences:

  • A search finds the wrong version. Ask for customer A's indexation clause, and the system may return customer B's because it is almost identical as text. The model writing the answer does not notice.
  • A similarity score is not a check. A score of 0.93 between contract text and an invoice line says they are about the same thing, not that the amount is right.
  • There is no fixed threshold. What counts as a "high" score differs by model and by type of text. Anyone who picks a threshold without testing on their own data is guessing.

The rule that follows is simple. You use embeddings to find. Amounts, quantities, dates and statuses you then retrieve from the system that manages them: the ERP, the billing system, the CRM. A well-designed system uses the embedding to find the right customer and the right contract, and then reads the indexation percentage from a field or from the source document itself.

Embeddings in revenue control

Where do embeddings fit in detecting revenue leakage? Mainly where information sits in text rather than in fields.

Contract terms that were never entered into a system. A framework agreement in PDF contains a clause about annual price adjustment. The ERP has no indexation. Embeddings help to find the passages about indexation, minimum purchase or rates for additional work across hundreds of contracts. Reading out the percentage and comparing it with the invoiced price is a separate, exact step. More on that pattern in revenue leakage from missed price indexation.

Additional work in free text. An engineer writes on the job sheet "extra pipework installed, customer agreed". Nobody adds a line for it in invoicing. Embeddings can surface job sheets with descriptions like that, so someone can check whether there is a matching invoice line.

Signals in email and notes. "Customer is considering another supplier" appears in a CRM note, in three different phrasings, spread over months. Semantic search finds them. Whether it is a real churn risk is determined by combining it with hard data: falling purchases, open tickets, a renewal date approaching.

The pattern is always the same. Embeddings produce a list of candidates. Exact data confirms or rejects. A person decides when in doubt. That is also how AI within Revenue Intelligence should work: each step does what it is good at.

Worked example: why a similarity score is not evidence

Worked example: suppose you have 600 customer contracts and you want to know which customers have an indexation clause that has not been applied. You have a system using embeddings search for passages about price adjustment.

  • The system finds 180 contracts with a passage closely resembling "prices are adjusted annually".
  • A manual sample shows that these include contracts stating that prices are fixed for the term. That sentence is about the same subject and therefore scores high.
  • Say 30 of the 180 contain such a fixed-price clause. Anyone who sends those 30 customers a corrective invoice without a second step is invoicing in breach of the contract.

The right approach: use the 180 as a candidate list, have the percentage, start date and exceptions read out verbatim for each contract, and compare those with the prices in billing. Only then do you have a finding in euros. The figures in this example are made up to show the mechanism, not as an average.

What you need to know before believing a vendor

A lot of software now advertises "AI search" or "semantic search". There are almost always embeddings underneath. Five questions to ask:

  1. Where are the embeddings created? An embedding model runs locally or at an external provider. In the second case your text, including names and amounts, goes to that provider. Ask where that happens and under which data processing agreement. See also how to protect personal data in AI systems.
  2. Can the text be recovered from an embedding? Not literally, but a vector is not anonymous either. Researchers have shown that much of the original text can sometimes be reconstructed. Treat a vector database with the same care as the source.
  3. Who may see what? If all documents sit in one vector database, the system still has to filter on permissions with every search. Otherwise a salesperson may see passages from HR files through an AI assistant.
  4. What happens when a document changes? A contract is amended or a customer cancels. The old embedding remains until someone replaces it. Ask how often and how that is updated.
  5. Where do figures come from? If an assistant says "customer X has an indexation of 3 percent", ask whether that figure comes from a field or from a text passage the model has interpreted. Only the first can be checked without opening the source.

Checklist for embeddings in your own environment

Use this list when assessing a tool with semantic search or an AI assistant over your documents:

  • Which sources are converted into embeddings, and which deliberately not?
  • Does the text sent to the embedding model contain personal data or contract amounts?
  • Is every search filtered on the user's permissions?
  • How soon after a change in the source is the embedding updated?
  • Do amounts, quantities and dates in answers come from a structured source, with a reference?
  • Is there a second, exact step before a retrieved passage leads to an action, such as an invoice or a customer email?
  • Has it been tested on your own documents how often the system returns the wrong but closely similar passage?

Frequently asked questions

Is an embedding the same as a language model?

No. An embedding model converts text into numbers and writes nothing. A language model writes text. In many systems they work together: the embedding model finds the right chunks, the language model phrases the answer.

Do I need a vector database?

Only if you have a lot of text you want to search by meaning. For structured data such as orders, invoices and CRM fields, an ordinary database with ordinary queries is better and more exact.

Can embeddings find revenue leakage?

They can help find places where leakage may sit, such as contract clauses or job sheets with additional work. Whether revenue is actually leaking only becomes clear when you compare the agreement found exactly with what was invoiced.

Are embeddings safe for confidential contracts?

That depends on where they are created and stored, and whether permissions are enforced at search time. A vector is not encryption. Treat it as sensitive data.

Do I need to recreate my embeddings if I switch model?

Yes. Embeddings from different models cannot be compared with each other. If you switch model, all documents have to be converted again.

Share this article
Knowledge base · AI technology

More in this cluster

All 22 topics in this cluster

More from AutoMaat

Rather know what this costs you specifically?

The Revenue Audit puts a euro amount on where your revenue leaks.

Plan the Revenue Audit