Why cohort analysis is needed
Aggregate metrics — average retention for the month, average LTV — hide important information: older cohorts with good retention mask the poor numbers of new ones. Cohort analysis lets you see the trend for each group separately.
The typical question cohort analysis answers: “Do buyers who came from paid social in Q3 retain better than those who came from paid search in Q4, or not?”
The structure of a cohort table
Cohort | M+0 | M+1 | M+2 | M+3
-------------|------|------|------|------
January | 100% | 22% | 15% | 11%
February | 100% | 25% | 17% | 13%
March | 100% | 28% | 20% | —
Rising retention from January to March can indicate an improving product, a change in traffic quality or the launch of a personalization programme. Cohort analysis makes that visible; an aggregate metric does not.
Types of cohort analysis in e-commerce
Acquisition cohort — grouping by the date of the first purchase or first visit. It shows how the quality of acquired buyers changes over time.
Behavioural cohort — grouping by an action: “users who made 3+ purchases in the first month”, “users who used the on-site search in their first session”. It lets you compare the LTV of different behavioural segments.
A/B cohort — comparing cohorts from the control and test groups on long-horizon metrics. This is a more powerful way of judging the effect of personalization than a short A/B test.
Cohort analysis and personalization
Tip: run a retention cohort analysis after every meaningful change on the site: a new recommendation block, a homepage redesign, the launch of a segmentation programme. The cohort after versus the cohort before is a simple way to measure the long-term effect.
Three signs of healthy cohorts:
– The retention curve flattens out rather than tending to zero
– New cohorts retain better than older ones, meaning the product is improving
– LTV is growing in the behavioural cohorts that receive personalization
Tools
Cohort analysis is available in GA4, Amplitude, Mixpanel and Tableau or Power BI (through manually built cohort tables on top of a data warehouse). To measure the impact of personalization precisely, it is better to work with warehouse data directly — you get far more freedom in defining the cohort and the period.