Batch versus streaming
The two main paradigms for exchanging data in an integration:
| Parameter | Batch import | Streaming |
|---|---|---|
| Freshness delay | Minutes to hours (set by the schedule) | Seconds |
| Load on the system | A peak at load time | Spread out |
| Suits | Historical data, the catalogue | Real-time behavioural events |
| Error handling | Simpler — the whole file is there to inspect | Harder — events pass through in a stream |
Most personalization platform integrations use a hybrid approach: batch for initialisation and catalogue refreshes, an event API for behavioural data.
A typical initial import
When a personalization platform is launched, the sequence usually looks like this:
1. Batch: product catalogue (every SKU with attributes, prices, stock)
2. Batch: order history for 6-12 months
3. Batch: view history for 1-3 months (where available)
4. Switchover: the Event API starts accepting events in real time
5. Scheduler: incremental catalogue batch - daily
Important: purchase history is a critical input for recommendation algorithms. Without historical data the system is cold and spends its first few weeks running on less precise strategies (popularity, trending).
File formats
CSV — the most common choice for catalogues:
sku,name,category,price,stock
ABC123,Nike running shoes,Sport,49.90,50
NDJSON — convenient for nested structures:
{"sku": "ABC123", "name": "Nike running shoes", "attributes": {"color": "white", "size": [40,41,42]}}
{"sku": "DEF456", "name": "Adidas backpack", "attributes": {"color": "black"}}
Common mistakes
- No idempotency. Re-running an import must not create duplicates. Implement upsert logic on a unique identifier (sku, order_id, user_id).
- Refreshing the catalogue too rarely. If prices and stock are updated once a day but change more often in reality, recommendations can show products that are not available.
- No monitoring of the load. A silent cron job that stopped working three days ago is a classic failure. Set alerts for the absence of a fresh import.