AutoMaat
Knowledge base· AI technology

Context windows explained for AI systems

What a context window is, why a larger window does not mean AI sees all your data, and how to prevent analyses that are silently cut short.

Ricardo Mastenbroek8 min read
Lees dit artikel in het Nederlands

A context window is the amount of text a language model can take into account at once when writing an answer: your question, the instructions, the documents supplied and the answer itself. Anything that does not fit does not exist for the model at that moment. A large window does not mean the model sees all your business data. It means the software around it has to choose what goes in, and that choice determines whether an analysis is right.

How does a context window work?

A language model has no memory between conversations and no access to your systems. With every call it receives a block of text and continues writing from it. That block is the context window. It is measured in tokens: pieces of text smaller than a word. A short common word is often one token; a long or unusual word, such as a compound term in a contract, is several. Numbers, dates and amounts are often split into several tokens.

That one block contains everything:

  • the software's system instructions ("you are an assistant that checks invoices"),
  • the description of the tools the model may use,
  • the earlier messages in the conversation,
  • the documents or data supplied,
  • and the answer the model still has to write.

The windows of modern models are large: hundreds of thousands of tokens, more for some models. That sounds like enough for everything. For a year of invoices, a CRM with thousands of deals and hundreds of contracts, it is not, and even if it fitted, putting it all in would not be a good idea.

Why is bigger not the same as better?

Attention is not evenly spread

Something being in the window does not mean the model pays equal attention to it. Models have got better at finding a detail in a long text, but with large volumes of data the chance increases that a line is skipped or linked wrongly. For a summary that often does not matter. For a check where one missed invoice line is the difference, it does.

Arithmetic does not improve with more text

A model with 4,000 invoice lines in its window does not add them up the way a database does. It estimates, summarises, rounds. Sometimes the result is right. You just do not know when. Totals, differences and counts should be calculated by code or a database, with the model translating the question and explaining the result. That distinction is also central to RAG vs database queries for business data.

Cost and speed

Every token in the window is processed on every call. Filling a large window therefore costs compute, time and money, again with every question. How that works is covered in inference explained.

Truncation happens silently

This is the most dangerous point. If the data does not fit, the software has to leave something out. Poorly built software simply cuts the list off: the first 500 orders go in, the rest do not. The model is not told and writes an answer as if it had seen everything. "All orders in Q3 have been invoiced" may then mean: all orders that fitted.

What does this mean for revenue analysis?

Suppose you ask an AI assistant: "Which deals won this year have no invoice?" There are three ways software can approach that.

Everything in the window. All deals and all invoices are supplied as text. The model compares them itself. This works with ten deals. With thousands of deals it becomes slow, expensive and unreliable, and you run the risk of silent truncation.

A selection in the window. The software first looks for the deals that seem relevant and supplies those. Better, but then the question is who decides what is relevant. A selection based on similarity (see embeddings explained for business software) misses precisely the deals that look different.

Calculating outside the window. The software runs a query that links deals and invoices on customer and amount, and supplies only the result: 14 deals without an invoice, with the details. The model explains, ranks and proposes an action. The window then contains little, but the right thing. This is the approach that scales and that you can recalculate.

The third approach is what a good Revenue Intelligence system with AI relies on. The model does not need to see everything. It needs to have the right question executed and explain the outcome well.

Context windows and agents

With an AI agent that takes several steps, the window fills up as it works. Every time the agent calls a tool, the result is added to the window. An agent that requests twenty invoices and looks up the order for each one has a full window after a while.

Software usually handles that by summarising or dropping older steps. That can lose information that was needed later. An agent that read at the start that customer X has different payment terms may have lost that twenty steps later. That is not a fault of the model but of how the task was split up. Smaller tasks with a clear outcome are more reliable than one long session. More on the tools agents work with in what is tool calling.

Worked example: what actually fits

Worked example: suppose you have 3,000 invoice lines a year. Each line contains customer name, invoice number, date, item, quantity, price and VAT. As text, that is easily dozens of tokens per line. Take 50 tokens as a rough assumption.

  • 3,000 lines times 50 tokens is 150,000 tokens.
  • Add the corresponding orders, roughly the same size, and you are at 300,000 tokens.
  • Add the contracts containing the price agreements, and you are well above what most models process in one go.

And that is one year, without CRM data, without credit notes, without additional work. The actual number of tokens depends on your data and the model. The point is the order of magnitude: the revenue data of a company with EUR 2M+ in revenue does not fit meaningfully in one window. Filtering and calculating outside the model is not an optimisation; it is a precondition.

How do you tell whether a system handles this well?

Ask of every AI tool that reports on your data:

  1. How much data did the model see for this answer? A good system can show it: 1,240 orders, 1,198 invoices, period January to September.
  2. What happens if it does not fit? Is the user warned, or is the data silently truncated?
  3. Are totals calculated or written? Ask for a figure and then ask how it was produced. If the answer is "the model worked it out", be careful.
  4. Can I open the source of a statement? Every claim about a customer or invoice should trace back to a record.
  5. How long are the agent sessions? An agent that works for hours in one window loses information along the way.

Checklist for your own use of AI on business data

Even if you simply use a chat assistant with an uploaded file, the same rules apply:

  • Do not upload an export of thousands of lines and ask for a total. Have the total calculated in Excel or your accounting software.
  • Ask the model to state how many lines it read. Compare that with the number in the file.
  • Break big questions down. First per customer, per month or per product group.
  • For every finding, ask for the invoice number or order number, so you can check it.
  • Be extra critical of "everything is correct". A model that did not see half the data will not find errors in it either.

Frequently asked questions

What is a token?

A token is the piece of text a language model calculates with. Often a word or part of a word. How many tokens a text costs depends on the language and the model. Text in languages other than English, and numbers, generally cost more tokens than the word count would suggest.

Does an AI model remember what I asked yesterday?

The model itself does not. Software can store earlier conversations or notes and put them back into the window with a new question. That feels like memory, but it is the software choosing what goes in.

Is a larger context window always better?

No. It gives room, but costs more and does not guarantee that the model uses everything equally well. For business data, the selection of what goes in matters more than the maximum size.

Can I load my whole CRM into an AI assistant?

Technically you can upload a lot, but the model will not calculate with it the way a database does. For questions about counts and amounts, you need an integration that runs queries and gives only the result to the model.

How do I notice that data has been truncated?

Often you do not, and that is the problem. So always ask how many records were included and compare that with the source. A system that cannot show this is unsuitable for control work.

Share this article
Knowledge base · AI technology

More in this cluster

All 22 topics in this cluster

More from AutoMaat

Rather know what this costs you specifically?

The Revenue Audit puts a euro amount on where your revenue leaks.

Plan the Revenue Audit