The structure of a neural network

A neural network is organised into layers. Data passes through them left to right — from the input
layer to the output layer:

  • Input layer — receives the features: user ID, item ID, session attributes, behavioural
    embeddings.
  • Hidden layers — apply linear transformations (multiplication by a weight matrix) and
    non-linear activations. This is where the network learns the features it actually needs.
  • Output layer — produces the prediction: a purchase probability, a relevance score, an
    embedding vector.
Input: [user_id, item_id, session_depth, hour_of_day, ...]
  → Dense(256, ReLU)
  → Dense(128, ReLU)
  → Dense(64, ReLU)
  → Dense(1, Sigmoid)  → P(purchase) = 0.73

How a neural network trains

Training is an iterative process. At every step:

  1. Forward pass: the data flows through the layers and the network produces a prediction.
  2. Error computation: the loss function compares the prediction with the true answer.
  3. Backpropagation: the gradient of the error is propagated backwards through the layers.
  4. Weight update: an optimiser (SGD, Adam) moves the parameters in the direction that reduces
    the error.

After millions of such steps the network settles on weights that minimise the error on the training
set.

Neural networks in e-commerce

Task Architecture Input
Recommendations Two-Tower, MLP User and item embeddings
Search ranking BERT / Transformer Text queries plus product attributes
Churn prediction MLP, LSTM Purchase sequence, visit frequency
Visual search CNN + CLIP Product images
AI assistant Transformer (LLM) Conversation context plus catalogue

Important: Neural networks need considerably more data to train than traditional methods such
as gradient boosting or collaborative filtering. Below roughly 100K–500K transactions a month they
often lose to simpler models because they overfit.

Neural networks versus traditional recommendation algorithms

Collaborative filtering (matrix factorization) works well on large user–item interaction datasets,
but cannot use side features such as session context, time of day or product attributes. Neural
networks combine every feature type inside one model, which is where they gain the advantage once a
rich feature space is available.