The core concepts of RL
Reinforcement learning is built around the interaction between an agent and an environment:
- Agent — the system making decisions (a recommendation engine, a conversational assistant).
- Environment — the context of the interaction (the shopper on the site, their session, their
history). - State — the current context: the page, the user’s history, the time of day.
- Action — the system’s choice: show product A or B, ask question X or Y.
- Reward — the feedback signal: a click = +0.1, a purchase = +1.0, an exit = 0.
- Policy — the action-selection strategy the agent is optimising.
The RL loop:
State → Agent → Action → Environment → Reward + New State
↑_______________|
Multi-armed bandit: RL in A/B testing
MAB is the most widespread application of RL in e-commerce. The name is an analogy with a slot
machine: the player faces several levers — the variations in an A/B test — each with an unknown
probability of paying out, that is, converting. The task is to maximise the total payout while still
exploring the options.
| MAB algorithm | Principle | Character |
|---|---|---|
| Epsilon-greedy | Exploit the best with probability 1-ε, explore at random with probability ε | Simple, suboptimal |
| UCB | Explore the variations with high uncertainty | Deterministic, good on low traffic |
| Thompson Sampling | Bayesian probability estimates with sampling | Fast convergence, often best in practice |
Tip: Thompson Sampling is the preferred choice in e-commerce — it concentrates traffic on the
winner quickly when the gap between variations is large, yet keeps exploring while uncertainty
remains. That reduces the conversion lost compared with a classic A/B test when variations are
unequal.
Exploration versus exploitation in practice
The RL dilemma in a recommendation context: show the shopper more of the product types they already
liked (exploitation), or try new categories (exploration)?
Exploitation only: the shopper sees the same things repeatedly, discovers new categories poorly,
and variety collapses.
Exploration only: irrelevant products get shown for the sake of learning, and conversion is lost
here and now.
In practice recommender systems use soft strategies: 80% to 90% of the slots from exploitation (the
relevant items), 10% to 20% from exploration (new categories chosen on affinity signals).
Deep RL
For complex systems — conversational assistants, marketing campaign control — deep RL is used: a
neural network approximates the value function (the Q-function) or the policy directly. That makes
it possible to handle rich states and delayed rewards, such as a purchase made three days after the
click rather than immediately. It demands considerably more data and computation than MAB.