Batch versus streaming

The two main paradigms for exchanging data in an integration:

Parameter Batch import Streaming
Freshness delay Minutes to hours (set by the schedule) Seconds
Load on the system A peak at load time Spread out
Suits Historical data, the catalogue Real-time behavioural events
Error handling Simpler — the whole file is there to inspect Harder — events pass through in a stream

Most personalization platform integrations use a hybrid approach: batch for initialisation and catalogue refreshes, an event API for behavioural data.

A typical initial import

When a personalization platform is launched, the sequence usually looks like this:

1. Batch: product catalogue (every SKU with attributes, prices, stock)
2. Batch: order history for 6-12 months
3. Batch: view history for 1-3 months (where available)
4. Switchover: the Event API starts accepting events in real time
5. Scheduler: incremental catalogue batch - daily

Important: purchase history is a critical input for recommendation algorithms. Without historical data the system is cold and spends its first few weeks running on less precise strategies (popularity, trending).

File formats

CSV — the most common choice for catalogues:

sku,name,category,price,stock
ABC123,Nike running shoes,Sport,49.90,50

NDJSON — convenient for nested structures:

{"sku": "ABC123", "name": "Nike running shoes", "attributes": {"color": "white", "size": [40,41,42]}}
{"sku": "DEF456", "name": "Adidas backpack", "attributes": {"color": "black"}}

Common mistakes

  • No idempotency. Re-running an import must not create duplicates. Implement upsert logic on a unique identifier (sku, order_id, user_id).
  • Refreshing the catalogue too rarely. If prices and stock are updated once a day but change more often in reality, recommendations can show products that are not available.
  • No monitoring of the load. A silent cron job that stopped working three days ago is a classic failure. Set alerts for the absence of a fresh import.