Four cycles running at the same time
Seasonality is not just “people buy fans in summer”. In e-commerce, at least four independent cycles
are layered on top of any data series, and confusing one for another is the source of most wrong
conclusions.
| Cycle | Period | How it shows up | What it distorts |
|---|---|---|---|
| Yearly | 12 months | Category seasons: tyres, air conditioners, back-to-school, gifts | Sales plans, quarter-on-quarter comparison, long experiments |
| Weekly | 7 days | Weekdays versus weekends, different demand mix and order value | Tests not run in whole weeks, weekly reports with a skew |
| Daily | 24 hours | Morning and evening traffic peaks, the overnight trough | Tests launched for a few hours; conclusions from a partial day |
| Event-driven | Calendar dates | Sales, holidays, pay days, the start of the school year | Comparison of adjacent weeks, extrapolating promo results to normal periods |
The event cycle is the least obvious and the most aggressive of the four. Pay days create a stable
two-peak structure across the month, and a week-long sale can produce a bigger amplitude than a
category’s entire yearly season.
How to separate seasonality from the effect of a change
The base mistake is the before-and-after comparison: measure two weeks before the release, two weeks
after, and book the difference as the effect. That difference contains everything at once — the
change, the season, advertising, assortment shifts and the weather.
Techniques, in ascending order of reliability:
- Year-on-year comparison. The same calendar period last year as the baseline. It works for
planning but not for evaluating changes: everything else moved between the two years as well. - A 7-day moving average. Smooths the weekly cycle and makes the trend visible. It does not
remove the yearly or event cycles. - Decomposing the series into trend, seasonal component and residual. Useful for demand
forecasting and purchasing, but the residual still contains every external factor except the one
under study. - A control group. Both groups live through the same season, on the
same days, with the same advertising. The difference between them is free of season by
construction — not approximately, but exactly.
No amount of mathematics rescues a before-and-after comparison. If a change was rolled out to
everyone at once, its effect cannot be separated from the seasonal and advertising background — only
the order of magnitude can be estimated. The one reliable decision is taken in advance: roll out
through an A/B test or keep a holdout group.
The effect on A/B tests
Seasonality breaks experiments in three ways, and all three are avoided by disciplined planning.
Incomplete weekly cycles. A test that runs for 3, 5 or 10 days gives a different balance of
weekdays and weekends in different stretches, and behaviour differs across those days — from the
traffic mix to the order value. The rule: test duration is a whole number
of full weeks, starting and finishing on the same day of the week.
A test inside an abnormal period. A result obtained on Black Friday describes the behaviour of
Black Friday shoppers. The traffic mix is different (more discount hunters), the motivation is
different and the sensitivity to urgency is different. That conclusion cannot be moved to an ordinary
Tuesday.
Rule for transferring conclusions:
test in a promo period -> conclusion applies to promo periods
test in a normal period -> conclusion applies to normal periods
test across a period boundary -> conclusion applies nowhere,
the result is mixed
Stopping early on a seasonal peak. The temptation to stop a test when a variation is “clearly
winning” is strongest on high-traffic days, because significance accumulates fast. But a spike that
coincides with a promotion easily produces a false positive. The same rules apply here as against
peeking: a duration and a sample size fixed in advance.
One more point: if a test has to run across a season boundary (catching the start of a sale, for
example), break the result down by period and check whether the direction of the effect holds in
each. Different directions signal that the effect depends on context and that there is no single
conclusion.
The effect on recommendations
Recommendation algorithms learn from history, and history by definition belongs to the previous
season. Without constraints, that leads to predictable artefacts:
- in May, New Year gift sets stay at the top of the recommendations, because their winter numbers were record-breaking;
- at the height of a season the algorithm is slow to lift new products that have no accumulated statistics yet — the classic cold start on seasonal assortment;
- once a season ends, recommendations keep dragging the departing category along for several weeks.
What teams do about it:
| Technique | What it delivers |
|---|---|
| Limiting the training window | The model stops treating last year’s season as a current signal |
| Higher weight on recent interactions | Fast reaction to a shift in demand inside the week |
| Scheduled rotation of strategies | Different strategies for seasonal, promo and normal periods |
| Filtering by availability and seasonal category | Products that are out of stock or out of season leave the output |
| Merchandising rules on top of the algorithm | Manual pinning of the seasonal top without rebuilding the model |
These settings have to be proven like any others: the seasonal configuration against the baseline in
the same period, on the same audience. Otherwise “the seasonal strategy worked” only ever means “the
season started”.
Checklist
- The yearly demand profile is known for every key category, not just for the store as a whole.
- The weekly and daily profiles are calculated separately — for traffic, conversion and order value.
- The calendar of event peaks (sales, holidays, pay days) is documented and used when planning releases and tests.
- No decision about the effect of a change is taken from a before-and-after comparison.
- Experiment duration is a whole number of weeks, starting and finishing on the same day of the week.
- Tests do not start a few days before a major sale and do not finish inside one.
- For long-lived tools, a holdout group is kept running across several seasons.
- In recommendations, history depth is limited, strategy rotation is configured and a stock filter is in place.
- Reports on tool performance always state the period and its character (normal or promotional).