Call · 15 min
AI systemsMattia Esposito2 September 20267 min read

A chatbot on your knowledge base. It answers with the source, or it doesn't answer.

The same fifteen questions, every day, about the same spec sheets and the same policies. This piece answers from your own documents, shows the source next to every answer, and when the answer isn't there, it says so.

In brief

Every answer shows its source. Next to it you see the document and the exact passage it came from, so checking takes a second rather than a phone call.

"I don't know" is a feature. Outside the agreed scope, the assistant stops and hands over to a person, because an assistant that always answers will eventually make something up.

Retrieval reduces made-up answers; it doesn't eliminate them. Three legal tools built exactly this way, and marketed as hallucination-free, got it wrong between 17% and 33% of the time, according to Stanford.

This page covers a single piece of the system. The other pieces, and how we choose which one to start with, are on the services page.

The problem is the fifteen questions that come back every day

Every business has a small set of questions that come in constantly, always the same ones, and the answers are already written down somewhere. The cost isn't any single answer; it's that someone looks for it in a different file every time, and stops doing something else in the meantime.

The second cost only shows up later: the person answering goes from memory, and two people's memories give two different answers. A simple question turns into a contradiction in front of a customer, and it's the kind of mistake nobody ever records.

A question asked fifteen times a day isn't a patience problem. It's a document nobody has put where it's needed yet.

What changes in practice

Recurring questions stop landing on a person's desk. The person asking gets the answer straight away, with the source next to it, and the person who used to answer gets their day back without those fifteen interruptions.

The second change is that there's only one answer. Everyone gets what's in the approved documents, not what someone remembers reading, and contradictions in front of customers disappear because there's a single source.

The third is information nobody has ever had: a list of the questions the assistant couldn't answer. It's the most honest way to find out which documents are missing, and it arrives without anyone having to ask.

Answering from documents reduces made-up answers, but doesn't eliminate them

This is where people buy badly, and there's evidence. Stanford's RegLab carried out the first preregistered evaluation of three legal research tools built on document retrieval and marketed as hallucination-free.

The result is stated plainly in the paper: “each hallucinate between 17% and 33% of the time”. The authors also note that hallucinations are reduced compared with a general-purpose assistant, which is exactly the point: reduced, not removed.

On a simpler task, the numbers are much better. The Vectara leaderboard, updated on 11 May 2026 across more than 7,700 documents, puts the best model at 1.8% of answers that aren't faithful to the text provided, with 99.5% of questions answered.

Taken together, the two measurements tell you something useful. Summarising a document you've been handed is reliable; picking the right document yourself and then answering is much less so. Almost all the build work goes into that second step.

You decide the scope, up front

During the analysis we decide what the assistant can answer, what it has to refuse, and what it has to pass to a person. It's a commercial decision, not a technical one, and it has to be made before a single line is written: an assistant with no stated scope will try to answer everything.

Type of searchWhat the assistant doesWhat you get
Approved informationit's in the documents

It answers straight away, showing the document and passage it took the answer from.

The same answer every time, which the person receiving it can check in a second.

Missing informationit's nowhere

It says it doesn't know, hands the question to a person, and adds it to the list of things to fill in.

No made-up answers, and a list showing which documents are really missing.

Anything that commits the businessprices, discounts, confirmations

It never answers on its own. It gathers the request with its context and queues it for the person who makes the decision.

No figure promised by a machine, and the person who decides gets the request with everything already gathered.

Off-topicnothing to do with it

It declines and steers the conversation back within scope, without chasing the question.

The assistant stays yours, and doesn't become a toy for people trying to throw it off course.

What the system does, step by step

Every question goes through the same five steps, and each one leaves a log entry with the question, the documents retrieved and the answer given. That log is how mistakes get found, instead of being pieced together from memory.

StepWhat happensWhat you get
The knowledge baseapproved documents

Spec sheets, terms, policies and past FAQs go into the knowledge base, each with the date it was last checked.

Answers come from documents someone has approved, and you know when.

Retrievalwhich passages are needed

When a question comes in, the system pulls up the relevant passages and gives them to the model, which answers only from those.

The answer rests on your own text, not on whatever the model happens to remember about the world.

Citationwhere it came from

The document and the exact passage appear next to the answer, clickable for anyone who wants to check.

Checking takes a second, and a mistake shows up before it becomes a problem.

Refusalwhen it doesn't know

If the retrieved documents don't contain the answer, the assistant says so and hands the question to a person.

The worst case becomes a delay, not wrong information delivered with confidence.

Testinga fixed set of questions

A set of real questions, each with the correct answer, is rerun every time the system changes.

Any decline in quality shows up immediately, instead of being discovered by a customer three months later.

The knowledge base is the product; the model is a replaceable part

An assistant is only as good as the documents behind it. If your terms of sale exist in three versions and nobody knows which is current, the assistant will repeat the wrong one with exactly the same confidence as the right one.

That's why the work starts with the documents, not the technology: you pick one good version for each topic, date it, and decide who keeps it up to date. It's the same principle as the set of facts behind sales copy and materials, and it's best to keep a single one for both.

Whether to write the information straight into the instructions or build a retrieval system depends on volume and how often things change. A few pages that rarely change need no extra set-up, and adding one would be a cost that buys nothing.

The assistant says it's an AI system

From 2 August 2026, Article 50 of Regulation (EU) 2024/1689 applies, requiring providers to ensure “that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system”.

The disclosure comes before the conversation and takes one line. What it actually means for a small business, with the scope and the four things to get in order, is on our page on Article 50 of the AI Act for Italian SMEs.

In practice it works in your favour commercially. Anyone who finds out on their own that they were talking to a machine stops trusting even the parts that were true, and that opening line removes the risk rather than adding one. The full reasoning is on our page about the principles we build by.

What goes out on its own, and what waits for a person

Anything that repeats information already approved goes out on its own: opening hours, availability, what's on a spec sheet, written terms, internal procedures. These were all decided in advance by a person, and the system simply passes them on, citing where they come from.

Anything that commits the business never goes out on its own. A price, a discount, a confirmation or a delivery promise is gathered with its context and queued for the person who makes the decision. Enquiries that arrive through several channels and need sorting are a separate piece, handled by Inbox AI, the first reply to every enquiry.

Same mechanism, different name in every trade

The mechanism is identical; the knowledge base isn't. It's worth looking at your own case, because that's where you see which document to sort out first.

SectorWhat goes in the knowledge baseWhere you see it
Food and agricultureand export

Spec sheets, allergens, pack sizes and units per case, delivery terms. The questions a buyer asks before asking for a price.

A typical day for a small food producer who exports

Restaurantsand bars

The current menu, allergens, opening hours, rules for groups and pets. Questions that pour in during the hour when nobody can answer.

A dinner service with the phone ringing unanswered

Hospitalityaccommodation and events

House rules, what's included, room capacities, check-in and check-out times, nearby services.

The run-up to the season for a guest accommodation business

What this piece doesn't do

It doesn't promise zero errors, and be wary of anyone who does: the Stanford study exists precisely because three big providers had made that promise. What gets built is a system where mistakes are visible, traceable to a source and fixed, with a fixed set of test questions that keeps quality in check over time.

It doesn't replace the person who answers. It takes away the repetitive questions and lets the ones that need judgement reach a person sooner, with more context.

It doesn't learn from conversations on its own. Anything that goes into the knowledge base does so because someone approved it, and a wrong answer given yesterday doesn't become tomorrow's truth.

Questions and answers

If it answers from our documents, can it still make things up?

Yes, less than before, but yes, and anyone promising otherwise is promising something nobody has proved. Stanford's RegLab evaluated three legal research tools built on exactly this approach and marketed as hallucination-free, and found they got it wrong between 17% and 33% of the time.

That's why the system shows the source document next to every answer, so checking takes a second, and why the scope of allowed questions is agreed in advance.

What happens when the answer isn't in the documents?

The assistant says it doesn't know and passes the question to a person, with a note of what was asked. That's a feature, not a fault: an assistant that always answers will eventually make something up, because it has no other way to answer.

Questions that fall outside the scope are collected in a list, and that list is the most honest way to find out which documents are really missing from the knowledge base.

Do we need a document retrieval system, or is a prompt enough?

It depends on how much knowledge there is to handle, and it's decided during the analysis using a simple rule. If the information fits on a few pages that rarely change, it goes straight into the assistant's instructions and no retrieval set-up is needed.

If there are dozens of documents that keep changing, you need a system that indexes them and pulls up the right ones when a question comes in. The choice depends on volume and how often things change, never on adding another technology to the quote.

Does the assistant have to say it's an AI system?

Yes. From 2 August 2026, Article 50 of Regulation (EU) 2024/1689 applies, requiring providers to ensure that people are informed they're interacting with an AI system. The disclosure comes before the conversation, not after, and it's a single line.

In practice it also works in your favour commercially, because anyone who finds out on their own that they were talking to a machine stops trusting even the parts that were true.

How do you measure whether it's working?

By how accurate its answers are against the actual knowledge base, measured on a fixed set of questions drawn up before we start. We write down the questions that genuinely come in, each with the correct answer from the documents, and rerun the same set every time the system changes.

The second number is the share of questions the assistant said it didn't know the answer to. It should stay steady, because if it suddenly drops, the assistant has started answering things it doesn't know.

Notes on sources

  1. The 17% and 33% figures come from the first preregistered evaluation of legal research tools built on document retrieval, by Stanford's RegLab, with the full text in the paper Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. It tests three US legal tools on questions of law, a far harder domain than questions about a spec sheet: the figure isn't a forecast for your case, it shows that the approach alone isn't enough.
  2. The 1.8% and 99.5% figures come from Vectara's public leaderboard, updated on 11 May 2026 across more than 7,700 documents. It measures how faithfully a model summarises a text it has been given, so the easier of the two tasks, and it's published by the company that sells the evaluation model: we cite it because the method and data are public and can be checked.
  3. The obligation to tell people they're talking to an AI system is in Article 50 of Regulation (EU) 2024/1689, published in the Official Journal of the European Union, applicable from 2 August 2026. What it covers for a small business is set out on our page on the AI Act for Italian SMEs. This page isn't legal advice: the exact scope of an obligation depends on the case, and it's a question for your own adviser.
  4. We don't publish any forecast of the share of enquiries an assistant will end up handling. The estimates going around are analyst projections about markets and future years, and they say little about an individual business with its own knowledge base. That share is measured on your own case, by counting a month's worth of real questions.
  5. This page doesn't report results achieved for a client, because this piece hasn't yet been delivered to a client. The tests mentioned are functional checks run in a test environment.
·The next step

Fifteen minutes, with your case in front of us.

For one week, note down every question you get asked more than three times, and next to each one write which file holds the answer. If there's no file for half of them, that's where the real work is. In fifteen minutes on the phone we'll look at it together and tell you where it makes sense to start, even if we never end up working together.

You'll speak to Mattia Esposito, who then builds the system: there's no salesperson in between. If you'd rather measure things yourself before talking, the Diagnostico (in Italian) is twenty questions and five minutes.