What feature engineering is
Feature engineering is the conversion of raw data into numeric representations a model can be
trained on. A machine learning model cannot work with the phrase last visit was the day before
yesterday, but it works perfectly with the number 2 — days since the last visit.
In e-commerce the raw material is click logs, transactions, product attributes and customer
profiles. Feature engineering turns those into informative variables:
Raw data → Features
─────────────────────────────────────────────
Click log → views per week, CTR by category
Purchase history → R (recency), F (frequency), M (monetary)
Product attributes → category (one-hot), brand, price bin
Session data → scroll depth, time on page
Types of transformation
Aggregating behavioural data
The most important class of features for recommendations is behavioural aggregates weighted by
recency. The event bought 14 days ago carries more weight than bought six months ago. The typical
pattern is rolling windows of 7, 30 and 90 days:
| Feature | Description |
|---|---|
| purchase_cnt_30d | Number of purchases in 30 days |
| avg_order_value_90d | Average order value over 90 days |
| days_since_last_visit | Recency of the last visit |
| top_category_share | Share of the top category in purchases |
Encoding categorical variables
Categorical attributes such as brand or category cannot be fed to a model directly. The main
approaches:
– One-hot encoding — for low-cardinality features (device type, gender)
– Target encoding — the mean of the target metric per category, suited to brands and categories
with thousands of values
– Embeddings — for entities with very high cardinality (product IDs, customers)
Feature engineering versus automatic approaches
Deep learning on unstructured data — text, images — extracts features by itself. On tabular,
structured data, manual feature engineering remains critical.
Tip: always inspect feature importance after training. In real e-commerce problems, 20% of the
features carry 80% of the predictive power — the rest add noise and slow inference down.
Common mistakes
- Data leakage: a feature that carries information from the future, for example a purchased flag
used at the moment of predicting a purchase - Very high cardinality with no encoding: user_id as a plain categorical feature without embeddings
- Ignoring temporal dynamics: averaging the whole history instead of weighting it by recency