AutoMaat
Knowledge base· AI technology

Inference explained

Inference is the moment an AI model answers. What it costs, where your data is at that moment and why the outcome is not always the same.

Ricardo Mastenbroek7 min read
Lees dit artikel in het Nederlands

Inference is the use of a trained AI model: you give it input, it calculates and returns an output. Training happens once and creates the model; inference happens with every question and is the moment your data goes into the model. For business software, inference determines three things you need to know: what each question costs, where your data is at that moment, and how predictable the outcome is.

What is the difference between training and inference?

An AI model is created through training. The model is adjusted using enormous numbers of examples until it predicts well. That is expensive, takes a long time and, for large models, is done by their maker. After training, the model's settings are fixed. Those settings are called weights or parameters.

Inference is everything that happens after that. Every time someone asks a question, has a document summarised or lets an agent take a step, the model runs a calculation with those fixed weights on the new input. The model learns nothing in the process. What you enter today does not change the model tomorrow, unless the provider later uses that data for new training. That is a contractual question, not a technical one, and you need to ask it.

That separation matters to avoid misunderstandings. "The AI learns from our data" is not true of inference. What can happen: software stores information and supplies it again with a later question. That looks like learning, but it sits in the software, not in the model. If you want to change the model itself, you end up at fine-tuning. The difference from supplying data is covered in fine-tuning vs RAG for business software.

How does inference work in a language model?

A language model writes an answer token by token. A token is a piece of text, often part of a word. At each step the model calculates which tokens are likely to come next, picks one, appends it to the text and calculates again. An answer of 300 words is therefore hundreds of calculation steps in a row, each over the whole context window. What that window is and why it matters is covered in context windows explained for AI systems.

Two things follow from this.

Longer input and longer answers cost more. Computing time grows with what goes in and what comes out. Providers therefore usually charge per token, often with a separate rate for input and output.

The answer is not always the same. There is often an element of chance in choosing the next token. A setting called temperature determines how much. At zero, the model always picks the most likely token, but even then answers are not always exactly identical in practice. For a summary that is no problem. For a check where you want the same invoice to get the same verdict twice, it is.

Where does inference run, and why does it matter?

At the moment of inference, your data has to be with the model. So that is the moment your customer names, amounts and contract terms reach a server. There are broadly three variants.

Variant Where the model runs What it means for your data
API from a model provider At the model provider Data goes to that provider; terms on storage, region and reuse are in their conditions
Model via a cloud platform In a cloud environment, often with a choice of region Data stays within that cloud environment and region, under the terms with that platform
Own hardware On servers you manage yourself Data does not leave your environment, but you manage everything yourself and the available models are more limited

None of the three is good or bad by definition. What counts is that you know which it is and that it matches your data processing agreement. Ask every vendor of AI software: which model do you use, where does the inference run, in which region, and is my data stored or used for training? See also how to protect personal data in AI systems.

What inference costs, and why that shapes your architecture

A single question to a language model costs little. It is different when software calls a model for every record. A system that runs every invoice line through a language model every night to judge whether it is correct makes thousands of calls a night. That is slow, expensive and mostly unnecessary: whether an invoice amount equals the order amount is a comparison a database does in a fraction of a second.

That is why well-built systems divide the work:

  • Rules and queries do the arithmetic: linking, adding up, comparing. That is exact, cheap and repeatable.
  • Smaller, specialised models handle tasks such as scoring anomalies or classifying text. See what is anomaly detection.
  • A large language model only comes in where language is needed: reading a free-text clause, explaining a finding, phrasing a proposal.

That order keeps inference limited to where it adds something. It also matches how AI within Revenue Intelligence should work: the model is one part of a chain, not the whole chain.

Worked example: per record or per exception

Worked example: suppose you have 40,000 invoice lines a year and you want them all checked.

  • Approach A: every line through a language model. 40,000 calls, each with the invoice line, order and contract information in the window. That is a lot of tokens, a lot of waiting time, and every outcome is a judgement by the model that you cannot reproduce exactly.
  • Approach B: rules first, then the model. A query compares all 40,000 lines with the order and price list. Say 600 lines deviate. Only those 600 go to a model that looks for an explanation in the contract or the notes. That is 1.5 percent of the calls.

Approach B is not only cheaper. It is also easier to check: the 39,400 lines that are correct have been compared exactly, and the 600 discrepancies each have an identifiable reason. What the calls cost in euros depends on the model, the provider and the amount of text. The difference in order of magnitude remains.

Latency: why some answers take a long time

The time between question and answer is called latency. With inference it depends on the size of the model, the length of the input, the length of the answer and how busy the servers are. An agent that takes ten steps waits ten times.

For an overnight check that makes no difference. For something a salesperson wants to look up during a customer call, it does. Good software makes that trade-off per task: fast and small where it has to be, large and thorough where it can be.

Checklist: questions about inference for your vendor

  1. Which models do you use, and for which step in the process?
  2. Where does the inference run, in which region, and with which party?
  3. Is my input stored by the model provider, and for how long?
  4. Is my data used to train models? Is that in the contract?
  5. Which calculations are done by rules or queries, and which by a model?
  6. If I run the same check twice, do I get the same outcome? If not, how is that handled?
  7. What happens if the model provider is unavailable: does the system stop, or fall back on something else?

Frequently asked questions

Does an AI model learn from my questions?

Not during inference. The model uses fixed weights. Only if the provider later uses your data for training can it influence a future version. Whether that is allowed is in the terms, and it should be excluded in your contract.

Why does AI sometimes give a different answer to the same question?

Because there is often an element of chance in choosing each word, and because small differences in the input or on the server can lead to a different outcome. For control work, software should therefore keep exact calculations outside the model.

Is inference on your own servers safer?

Your data does not leave your environment, which is an advantage. But you are then responsible for security, updates and availability yourself, and you have less choice of models. It is only safer if you manage all that well.

What does inference cost?

That depends on the model, the provider and how much text goes in and comes out. There is no fixed price per question. What you can control is how often a model is called and with how much text.

Share this article
Knowledge base · AI technology

More in this cluster

All 22 topics in this cluster

More from AutoMaat

Rather know what this costs you specifically?

The Revenue Audit puts a euro amount on where your revenue leaks.

Plan the Revenue Audit