How Chatbots Fail, and What a Useful One Looks Like
Comprehension is usually not the problem
Public discussion about chatbots tends to fixate on whether the underlying model is capable enough. In our experience at SmartEdge IT Solutions the model is rarely the binding constraint. Systems built on capable models fail for ordinary operational reasons: the answer was confidently wrong, the bot addressed a question it had no record of, or the visitor could not reach a person when they needed one.

Those are design failures rather than model failures. They are also fixable, but only if you accept at the outset that a chatbot’s job is not to sound knowledgeable. It is to be useful on a bounded set of questions and honest everywhere else.
The failure everyone describes as “the bot made things up” is usually the bot doing what it was built to do — generating plausible text — in a situation where nobody gave it a way to check. A model given your refund policy will paraphrase it well. A model given nothing will fill the gap, because that is the only option available to it.
Breadth without depth
The standard demonstration covers a lot of ground. Ask about delivery, returns, opening hours, service tiers, escalation to a person, and the bot handles all of them fluently. What the demonstration cannot show you is the tenth follow-up question, the one where the visitor’s situation turns specific.

Real queries are rarely the frequently asked question. Somebody who has read the returns page and still has a question is, by definition, in the part of the process that page does not cover. Chatbots tend to be strongest exactly where the need is weakest and weakest where the need is real.
There is a second consequence. If the bot covers the easy half of the volume perfectly, visitors who would have rung you now type instead, and the easy half consumes capacity that does not free anything up. The value of a bot is not in handling the question the documentation already answers. It is in the questions it can answer despite the documentation.
A useful diagnostic: take the last hundred genuine contacts your team handled and sort them by whether the answer already existed in your published documentation. Then look at the remainder. That remainder is what a chatbot on your site would actually be handling, and it is a much narrower and more interesting list than the FAQ.
It is usually a list dominated by a handful of recurring situations — a damaged order, a duplicated charge, an exception to standard policy, a request needing a judgement call. Those are worth handling. The long tail around them is where chat goes wrong.
Deflection is not resolution
Chatbot platforms tend to report deflection rate: the proportion of conversations that never reach a human. It is an easy number to improve and a poor one to rely on, because a bot that frustrates people into going elsewhere posts an equally good number. The people leaving are not recorded anywhere.
The figure worth watching is resolution: conversations that ended because the visitor got what they needed, however that happened. It is harder to measure because you have to sample transcripts and read them, and because it requires asking the customer something at the end rather than inferring it from behaviour.
Where a chatbot sits in the journey also determines what it should be measured on. An assistant on a pricing page is doing a different job from an agent handling a failed delivery. Conflating them produces a single metric that means nothing and a business case that cannot be evaluated.
Decide the job first, in one sentence, and refuse a second one later. The feature request that arrives in month four — “can it also handle refunds?” — is where most deployments go wrong.
The handoff that loses the customer
Most frustrating chatbot experiences share a structure: the bot understands roughly, produces something nearly useful, and then the transfer to a human drops the context. The visitor restarts. The customer repeats themselves, which is the moment trust goes.

A good handoff carries four things: what the visitor asked, what has already been attempted, which account or order is involved, and what the bot believed the answer to be. That last one matters most and is usually missing. Without it the human has to reconstruct the bot’s reasoning, and may unknowingly contradict it.
The transfer has to be genuinely available. Not “leave a message and someone will reply within one business day” dressed up as a live handover. If that is genuinely all you can offer, say so on the first screen rather than after six messages.
There is a pattern worth copying here: let the visitor skip the bot at any point, from anywhere, without having to ask in the bot’s own language. A permanent “talk to a person” control beside the input box rather than a phrase the bot has to recognise.
Sometimes the answer is better search
Before building a conversational layer, test whether the real problem is retrieval. Plenty of sites have a help page that answers the question perfectly and a search box that cannot find it. Adding a chatbot on top of that hides the problem rather than solving it.

The test is cheap. Take fifty real questions and check whether the answer already exists somewhere on the site. If it does and visitors are not finding it, the fix sits in navigation, titles and internal linking, and it helps every other channel as well. That work is usually a matter of search engine optimisation applied to structure rather than to keywords, and it tends to be far cheaper than a conversational layer.
Where to put the effort also depends on how questions arrive. Somebody typing into a box at eleven at night while finishing a shift and somebody phoning at ten in the morning are in different states and want different things. Conversational interfaces suit the second case better than the first, where a page that answers in three lines is better. If the answer is genuinely not on the site at all, that is a content gap worth filling first — a chatbot will not fill it, it will interpolate around it.
Where the chatbot genuinely earns its place is where the question cannot be answered with a static page at all: something depending on the visitor’s own record, or on a live status. That is a different product, and it needs access to real data rather than a document collection. Where it also needs to take actions rather than only answer, the work crosses over from chat into AI agent development, and the risk profile changes with it.
Narrow scope beats a clever model
Every chatbot project that has worked for us has a short, specific description of what it covers, written somewhere visible to the team. “Delivery tracking, returns and damaged-item claims for orders under thirty days” is a scope. “Customer support” is not.
Narrow scope has three practical consequences. The content to load becomes finite, so retrieval quality can be made good rather than merely acceptable. The failure modes become enumerable, so they can be tested. And the team can tell visitors what the thing actually does, which removes most of the frustration that gets blamed on the technology.
It also makes stopping possible. If the bot covers two topics well and a third badly, the third can be removed without the whole thing collapsing. Broad bots have no removable part.
Model choice comes after all of this. A good retrieval setup on a capable but unremarkable model will outperform a weak setup on a larger one, and teams routinely spend the attention they should have spent on the content on the model parameter instead.
SmartEdge IT Solutions has shipped chatbots on modest models where the content was genuinely well organised, and stalled on expensive ones where it was not. That is not a general law, but it has held across enough projects to be worth stating plainly.
What the transcript is actually for
Conversation logs are the most useful artefact a chatbot produces and the one most frequently left unexamined. Not because nobody wants the data, but because nobody has decided who reads it and what they are permitted to change.

The weekly review is short and specific: read the logs where the bot gave no answer, the logs where the visitor repeated themselves, and the logs where a human took over. Twenty minutes, three questions each time.
- Was the correct answer present in the content the bot could reach? If it was, this is a retrieval or phrasing problem. If it was not, this is a content gap.
- Did the bot state a limitation, or did it improvise? Improvised answers cluster in a small number of topics and are worth writing out of the source material.
- Did the visitor ask the same thing twice? Repetition is the clearest available signal that the first answer failed.
Each review should produce at most one or two changes, written into the same place as the content itself. A chatbot receiving twenty edits a week is telling you its content model is wrong.
Recording which answers were corrected, and which were accepted without change, builds the evaluation set for later work. Without it, every improvement is a guess and you cannot show whether the last change helped.
Our AI chatbot development work treats the transcript review as part of the build rather than an operational afterthought, because the review is where the quality actually comes from.
Measuring a chatbot honestly
Four measures, taken together, tell you more than any dashboard the supplier provides.

Resolution by topic. Not an overall figure. Some topics will be genuinely handled and others will be a mess, and an average hides exactly the thing you need to act on.
Correction rate. The share of answers a human changed substantially or discarded. Trending over time, it is the closest thing to a quality signal that does not require reading transcripts.
Repeat contact. Visitors who return about the same issue within a short window. This catches failures that produce no complaint because the person simply gave up or found another route.
Where conversations stop. Exit questions and abandonment points cluster, and one cluster usually represents a question the content does not cover.
And one counter-metric: what the bot costs per resolved conversation, including review time. If the review is heavy enough that a person would have been faster, the honest conclusion is that this is a triage layer rather than a resolution layer. That is a perfectly good thing for it to be, as long as it is not costed as though it resolves.
None of this needs a large model or an elaborate platform. It needs a narrow scope, honest content, a real route to a person and somebody paid to read the logs. That is the difference between a chatbot people use and one they learn to type past.
