The structure of a neural network
A neural network is organised into layers. Data passes through them left to right — from the input
layer to the output layer:
- Input layer — receives the features: user ID, item ID, session attributes, behavioural
embeddings. - Hidden layers — apply linear transformations (multiplication by a weight matrix) and
non-linear activations. This is where the network learns the features it actually needs. - Output layer — produces the prediction: a purchase probability, a relevance score, an
embedding vector.
Input: [user_id, item_id, session_depth, hour_of_day, ...]
→ Dense(256, ReLU)
→ Dense(128, ReLU)
→ Dense(64, ReLU)
→ Dense(1, Sigmoid) → P(purchase) = 0.73
How a neural network trains
Training is an iterative process. At every step:
- Forward pass: the data flows through the layers and the network produces a prediction.
- Error computation: the loss function compares the prediction with the true answer.
- Backpropagation: the gradient of the error is propagated backwards through the layers.
- Weight update: an optimiser (SGD, Adam) moves the parameters in the direction that reduces
the error.
After millions of such steps the network settles on weights that minimise the error on the training
set.
Neural networks in e-commerce
| Task | Architecture | Input |
|---|---|---|
| Recommendations | Two-Tower, MLP | User and item embeddings |
| Search ranking | BERT / Transformer | Text queries plus product attributes |
| Churn prediction | MLP, LSTM | Purchase sequence, visit frequency |
| Visual search | CNN + CLIP | Product images |
| AI assistant | Transformer (LLM) | Conversation context plus catalogue |
Important: Neural networks need considerably more data to train than traditional methods such
as gradient boosting or collaborative filtering. Below roughly 100K–500K transactions a month they
often lose to simpler models because they overfit.
Neural networks versus traditional recommendation algorithms
Collaborative filtering (matrix factorization) works well on large user–item interaction datasets,
but cannot use side features such as session context, time of day or product attributes. Neural
networks combine every feature type inside one model, which is where they gain the advantage once a
rich feature space is available.