What unsupervised learning is for
In most e-commerce problems the data is not labelled: there are millions of events — views,
add-to-basket actions, purchases — but no explicit answer to “which segment is this shopper in?” or
“is this transaction anomalous?”. Unsupervised learning finds the hidden structure in such data by
itself.
The three main tasks
Clustering — splitting objects into groups with high internal similarity. Shoppers with similar
behaviour land in the same cluster. The result is a set of audience segments that can each be given
a different strategy.
Dimensionality reduction — compressing high-dimensional data, such as a purchase vector across
50,000 products, into a compact representation of 10 to 50 features. The algorithms: PCA, t-SNE,
UMAP, autoencoders. Used for visualising segments and for preprocessing before training.
Anomaly detection — finding the points that do not fit the overall structure. In e-commerce:
fraudulent orders, review manipulation, bot traffic.
Use in audience segmentation
Clustering on RFM (recency, frequency, monetary) is the classic example. K-means — or DBSCAN when
the clusters have an irregular shape — divides shoppers into segments with no rules set in advance:
Cluster 1: high F, high M → VIP customers
Cluster 2: recent R, low F → new customers
Cluster 3: distant R, formerly high F → churn risk
The output is a set of segments usable for content personalization, triggered communications and
differentiated offers.
Tip: clustering results have to be interpreted by hand. The algorithm will find the groups —
but naming them (“loyal”, “churning”, “new”) is the analyst’s job. Without a business reading, the
clusters stay just numbers.