Why this matters in e-commerce

In a general chatbot a hallucination is annoying. In an AI shopping assistant it hits the business
directly:

  • A shopper learns a specification that does not exist, buys, is disappointed → a return plus lost
    trust
  • The assistant states a wrong price confidently → a conflict at checkout
  • The bot invents compatibility, say a charger for a laptop → an incompatible order is placed

Causes and mechanics

An LLM does not know facts in the ordinary sense — it holds patterns from a training corpus. When
generating, it follows the most likely continuation rather than a verified fact.

Hallucinations cluster around three situations:

  1. Questions about very specific, rare or recent data (a new SKU, a niche product)
  2. Requests for numeric facts (prices, specifications, dimensions)
  3. Conflicts between several sources in the training data

Mitigations

RAG

The most effective approach for e-commerce. Before generating, the system retrieves relevant
documents from a vector database — catalogue, FAQ, category descriptions — and adds them to the
context. The model answers from those documents rather than from memory.

Prompt engineering

System prompt instructions reduce the rate: answer only from the provided catalogue; if the
information is absent, say so rather than inventing it.

Output verification

Post-processing: a separate model or deterministic logic checks that the SKUs and prices in the
answer really exist in the catalogue. References that do not resolve are removed or replaced.

Important: RAG does not eliminate hallucination entirely — when the context lacks the needed
information the model can still fill the gap. A correct architecture handles the not found case
explicitly and says so to the shopper.