How clustering works
Clustering is an unsupervised learning task: the algorithm finds structure in the data on its own,
with no rules set in advance. The single criterion is that objects inside a cluster should resemble
each other, while objects from different clusters should differ.
The process, using k-means as the example:
- K, the number of clusters, is chosen
- K centroids are picked at random
- Every object is assigned to its nearest centroid
- Centroids are recomputed as the mean point of their cluster
- Steps 3 and 4 repeat until convergence
Clustering algorithms
| Algorithm | How it works | When to use it |
|---|---|---|
| K-means | Splits by distance to centroids | Customer segmentation by RFM, AOV |
| DBSCAN | Groups dense regions | Anomaly detection, outlier segments |
| Hierarchical | Builds a tree of nested clusters | Assortment taxonomy, category analysis |
| Gaussian Mixture | Probabilistic cluster membership | Soft segmentation: a customer can belong to several clusters |
Applications in e-commerce
Customer clustering surfaces behavioural segments without hand-written rules:
Cluster 1: high LTV, buys rarely, large basket — "considered buyers"
Cluster 2: frequent purchases, low AOV — "browsers"
Cluster 3: bought once, disappeared — "one-off buyers"
Cluster 4: only buys on promotion — "discount hunters"
Each cluster needs its own personalization strategy.
Product clustering addresses cold start: a new SKU with no view history is assigned to the
nearest product cluster by its attributes and starts receiving traffic from recommendations.
Tip: do not read clusters mechanically. K-means produces mathematical groups; their business
meaning is decoded by analysts. Always validate a cluster by asking what these 50,000 customers
actually have in common.
Limitations
- K-means is sensitive to feature scale — normalise the data before clustering
- Clusters are unstable under random initialisation — use k-means++ or several runs
- The optimal K is rarely obvious; the elbow method gives a hint, not an answer
- Clusters are a snapshot in time; behaviour shifts, so clusters have to be refreshed