A customer asks your support bot whether they can switch their billing to a different currency. It says yes. Cheerfully. With steps.

You cannot switch currency. There has never been a way to switch currency. But nobody wrote that down anywhere, and your billing article happily explains how to change your plan, your card and your billing address, so the bot did the reasonable thing with what it had.

Then the screenshot lands in Slack, and someone asks the question we hear more than any other: why is it making things up?

It isn't making things up

Here's what actually happens when that question comes in. The system searches your help center for the passages that look closest to the question. It hands those few hundred words to the model. The model writes an answer from what it was handed.

That's the whole trick. It isn't reasoning about your product. It has never seen your product. It's summarising a handful of paragraphs that someone on your team wrote, possibly two years ago, possibly in a hurry.

Which means the answer your customer gets is decided by three things, and not one of them is the model.

Does the passage survive on its own?

Retrieval pulls pieces of pages, not pages. A section that opens with "once you've completed the steps above" arrives with no steps above it. The model either throws it away or, worse, uses it anyway, minus the condition that made it true.

Try this on your own docs and it gets uncomfortable fast. Open any article, pick a section in the middle, cover everything above and below, and read only that. Does it still answer something? Could you tell which product it belongs to?

Does anything in your docs say no?

This is the currency problem, and it's the one that catches everyone.

If your article explains at length what a feature does, and never states what it can't do, then nothing in the retrieved text rules anything out. Asked "can I do X", the model has capability language on one side and silence on the other. It fills the silence.

That is not a hallucination in any exotic sense. It's a fair inference from a document that never mentioned the limit. The fix isn't a better model. It's a sentence.

Do your sources agree with each other?

When your help center, your pricing page and a blog post from 2024 say three different things, the model does what a cautious person would do: it hedges, or it picks whichever version sounds most confident. That is often not the one you'd have chosen.

The good news, and the annoying news

The good news is that bot quality is mostly a documentation project, and the work is all stuff you already know how to do. Make sections stand on their own. Write down the limits. Stop contradicting yourself. Every one of those makes the article better for the human reading it too, so none of it is wasted if you change vendors next year.

The annoying news is that you can't prompt your way out of it. A better system prompt cannot invent a limit that nobody ever wrote down.

If you want to find out what yours would get wrong before your customers do, start with the ten free prompts. One article, about five minutes, and you will have a list.