How clustering works

Clustering is an unsupervised learning task: the algorithm finds structure in the data on its own,
with no rules set in advance. The single criterion is that objects inside a cluster should resemble
each other, while objects from different clusters should differ.

The process, using k-means as the example:

  1. K, the number of clusters, is chosen
  2. K centroids are picked at random
  3. Every object is assigned to its nearest centroid
  4. Centroids are recomputed as the mean point of their cluster
  5. Steps 3 and 4 repeat until convergence

Clustering algorithms

Algorithm How it works When to use it
K-means Splits by distance to centroids Customer segmentation by RFM, AOV
DBSCAN Groups dense regions Anomaly detection, outlier segments
Hierarchical Builds a tree of nested clusters Assortment taxonomy, category analysis
Gaussian Mixture Probabilistic cluster membership Soft segmentation: a customer can belong to several clusters

Applications in e-commerce

Customer clustering surfaces behavioural segments without hand-written rules:

Cluster 1: high LTV, buys rarely, large basket   — "considered buyers"
Cluster 2: frequent purchases, low AOV           — "browsers"
Cluster 3: bought once, disappeared              — "one-off buyers"
Cluster 4: only buys on promotion                — "discount hunters"

Each cluster needs its own personalization strategy.

Product clustering addresses cold start: a new SKU with no view history is assigned to the
nearest product cluster by its attributes and starts receiving traffic from recommendations.

Tip: do not read clusters mechanically. K-means produces mathematical groups; their business
meaning is decoded by analysts. Always validate a cluster by asking what these 50,000 customers
actually have in common.

Limitations

  • K-means is sensitive to feature scale — normalise the data before clustering
  • Clusters are unstable under random initialisation — use k-means++ or several runs
  • The optimal K is rarely obvious; the elbow method gives a hint, not an answer
  • Clusters are a snapshot in time; behaviour shifts, so clusters have to be refreshed