Open your bot's logs and you'll naturally go to the bad ones. The escalations, the thumbs-down, the conversation where someone typed "this is useless" in capitals.
Those are the cheap failures. The customer knew they were being failed, so they escalated, and a human sorted it out.
The expensive ones look like this: question asked, answer given, conversation ended, no escalation, no rating. By every dashboard you have, that's a success.
A successful-looking conversation has two possible meanings
Either the bot answered correctly and the customer went away happy, or the bot answered confidently and wrongly and the customer went away believing it.
Those are indistinguishable from the outside. Same length, same shape, same lack of complaint. The second one only surfaces later, as a support ticket about something else, a refund request, or a customer who quietly stops using a feature that they now believe doesn't work.
So the metric everyone reports — escalation rate, or thumbs-up rate — is measuring how often customers noticed, not how often the bot was right.
What to actually do on a Tuesday morning
Pull twenty conversations at random from the ones that ended without escalation. Not the worst twenty. Random ones.
For each, read only the question and the answer, and check the answer against what's actually true. Not against your documentation — against reality. Your documentation is what produced the answer, so checking one against the other tells you nothing.
Twenty is enough. If three of them are wrong, you have a fifteen per cent silent error rate and a very clear afternoon ahead of you.
Then ask why, not what
When you find a wrong answer, resist fixing the sentence. Find the passage it came from. Nine times out of ten the passage is fine in context and useless out of it, or the limit it needed was never written down anywhere.
That tells you which article to fix, which is worth far more than correcting one conversation nobody will have again.
Make it a habit, not an audit
Twenty conversations a week takes about half an hour and will tell you more about your help center than any analytics dashboard. It's also the only way to catch the failure mode where everyone's numbers look great and your customers are being confidently misled.
Twenty conversations, once a week, chosen at random. It is the only routine we know of that catches the failure where every number looks good and your customers are being confidently misled.

