What a chatbot is
A chatbot is a program that holds a dialogue with a user and completes their task: answering a
question, finding a product, performing an operation in the company’s systems. The interface can be
a widget on the site, a screen in a mobile app or a messenger.
The term covers technologically very different things, from a tree of buttons to a model with
catalogue access. Because of that, “we have a chatbot” on its own says nothing about capability or
about the cost of keeping it running.
Three generations of bots
| Generation | How it works | Flexibility | Cost of upkeep | Main risk | Where it fits |
|---|---|---|---|---|---|
| Scripted | A tree of buttons and fixed transitions | Low: only the branches provided | Grows linearly with the number of scenarios | The dead end of “I did not understand” | Narrow repeated tasks: order status, opening hours |
| NLU bots | Classify the phrase into an intent, then run a script | Medium: free wording, fixed answer | Dataset labelling and intent list upkeep | A misclassification sends the user into the wrong script | Support with heavy volumes of routine requests |
| LLM assistants on RAG | The answer is generated from connected sources | High: questions nobody described in advance | Lower on scenarios, higher on infrastructure and quality control | Invented facts — price, availability, terms | Product selection, advice in complex categories |
Moving to a new generation does not retire the previous one. A working setup usually contains all
three: quick button answers for frequent requests, intent recognition for routing, and a generative
layer — an LLM in a RAG pattern — for open dialogue.
Why e-commerce is not decided by conversational skill
The phrasing quality of modern models stopped being the bottleneck a long time ago. The difference
between a useful and a useless bot in an online store runs along two lines.
Catalogue access. A bot that cannot see live prices, stock and specifications is physically
unable to solve the user’s main task — choosing a specific product. It gives general advice the
user has already read in reviews. Technically this is solved either with a RAG pattern over the
product feed or by calling the store’s own functions
(function calling) for price, availability and order status.
Access to the user’s context. Viewed products, basket contents, order history, size and brand
preferences change the answer far more than any tone setting does. Without that context the bot
starts every dialogue from zero and asks questions the store’s own data already answers.
User request
-> intent and selection parameters
-> catalogue search (feed: price, stock, attributes)
-> user context (views, basket, history)
-> answer + product cards + next step
This is also what separates an advisory bot from a support bot: the first works in the logic of
conversational commerce and leads to a purchase, the second
takes load off the agents.
How to measure it
The most common mistake is reporting dialogue and message counts. Those numbers rise both when the
bot helps and when the user rephrases the same question three times.
| Metric | What it shows | How to calculate it |
|---|---|---|
| Resolved dialogue share | Usefulness without an agent | Dialogues completed without escalation / all dialogues |
| Dialogue-to-basket conversion | Influence on the path to purchase | Dialogues with an add-to-basket / all dialogues |
| Revenue per dialogue | Economic contribution | Attributed revenue / number of dialogues |
| Difference against a control group | Real uplift rather than self-selection | Conversion of sessions with the bot vs without, in a test |
| Escalation share | Load on support | Handed-over dialogues / all dialogues |
Comparing users who opened the bot with users who did not is invalid: people with stronger purchase
intent enter a dialogue more often. The effect can only be isolated with an
A/B test in which part of the traffic never sees the bot.
A reference point from AI Shopping Assistant pilots: +33%
revenue per visit, +9% conversion and +20% average order value; the earlier quiz-format version gave
+14% conversion to basket.
Risks and guardrails
- Hallucinated price and availability. The most expensive class of error: the user receives a
promise the store will not keep. The rule is that every factual value comes from the store’s
systems, not from the model. - Answers outside its competence. Legal, medical and financial phrasing should be explicitly
banned in the prompt and filtered on output. - No human in the loop. Complaints, returns and warranty cases go to an agent. A
human-in-the-loop design is needed not only for quality but as insurance when the model fails. - Tone and pressure. A bot that pushes hard towards a more expensive product erodes trust faster
than it lifts average order value. - Logs nobody reads. Dialogues should not only be stored but sampled and read regularly: that is
where the real scenarios show up, along with the places the bot drifts off course.
A pre-launch checklist
- Three to five tasks the bot closes at launch are defined — from support and search logs, not from
hypotheses. - The product feed with prices and stock is connected, and its refresh frequency has been checked.
- Factual data (price, availability, order status) comes from function calls, not from generation.
- Escalation to an agent is configured, along with a fallback for when the model is unavailable.
- Metrics are set before launch: resolved dialogue share, conversion to basket, revenue per
dialogue. - A control group is carved out, otherwise the uplift will be indistinguishable from the
self-selection of interested users. - A regular log review is assigned — weekly at the start, less often later.