What feature engineering is

Feature engineering is the conversion of raw data into numeric representations a model can be
trained on. A machine learning model cannot work with the phrase last visit was the day before
yesterday, but it works perfectly with the number 2 — days since the last visit.

In e-commerce the raw material is click logs, transactions, product attributes and customer
profiles. Feature engineering turns those into informative variables:

Raw data              → Features
─────────────────────────────────────────────
Click log             → views per week, CTR by category
Purchase history      → R (recency), F (frequency), M (monetary)
Product attributes    → category (one-hot), brand, price bin
Session data          → scroll depth, time on page

Types of transformation

Aggregating behavioural data

The most important class of features for recommendations is behavioural aggregates weighted by
recency. The event bought 14 days ago carries more weight than bought six months ago. The typical
pattern is rolling windows of 7, 30 and 90 days:

Feature Description
purchase_cnt_30d Number of purchases in 30 days
avg_order_value_90d Average order value over 90 days
days_since_last_visit Recency of the last visit
top_category_share Share of the top category in purchases

Encoding categorical variables

Categorical attributes such as brand or category cannot be fed to a model directly. The main
approaches:
– One-hot encoding — for low-cardinality features (device type, gender)
– Target encoding — the mean of the target metric per category, suited to brands and categories
with thousands of values
– Embeddings — for entities with very high cardinality (product IDs, customers)

Feature engineering versus automatic approaches

Deep learning on unstructured data — text, images — extracts features by itself. On tabular,
structured data, manual feature engineering remains critical.

Tip: always inspect feature importance after training. In real e-commerce problems, 20% of the
features carry 80% of the predictive power — the rest add noise and slow inference down.

Common mistakes

  • Data leakage: a feature that carries information from the future, for example a purchased flag
    used at the moment of predicting a purchase
  • Very high cardinality with no encoding: user_id as a plain categorical feature without embeddings
  • Ignoring temporal dynamics: averaging the whole history instead of weighting it by recency