What rate limiting is for
Rate limiting is a policy that caps how many requests a client may make in a given period. The goals:
- Protection from overload — preventing service degradation when traffic spikes.
- Fair resource sharing — one client cannot consume all the capacity.
- Protection from abuse — denial-of-service attacks, scraping, credential stuffing.
When the limit is exceeded, the API answers:
HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1716300000
Rate limiting algorithms
| Algorithm | Principle | Characteristic |
|---|---|---|
| Fixed window | A counter over a fixed period (a minute, an hour) | Bursts at the window boundary |
| Sliding window | A rolling window with no fixed boundaries | More precise, no boundary spikes |
| Token bucket | A bucket refilled at a set rate | Tolerates bursts up to the bucket size |
| Leaky bucket | A queue draining at a fixed rate | Strictly even processing |
Rate limiting in e-commerce integrations
An online store running a personalization platform calls the API on every page view: the product page requests recommendations, the listing requests ranking, the home page requests banners. At 100K unique visitors a day that is roughly 3 to 5 million API calls in 24 hours, with peaks of several thousand requests per second.
Practical rules for integrations:
- Cache client-side. Recommendations for a given user identifier rarely change within one session — a cache with a 5 to 15 minute TTL cuts request volume threefold to fivefold.
- Use batch requests. Where the API supports it, fetch recommendations for several widgets in one call.
- Circuit breaker. After a run of 429 errors, switch automatically to a fallback — popular products from a local cache.
- Exponential backoff. Never retry immediately after a 429.
import time
def api_request_with_retry(url, max_retries=3):
for attempt in range(max_retries):
response = requests.get(url)
if response.status_code == 429:
wait = 2 ** attempt # 1, 2, 4 seconds
time.sleep(wait)
continue
return response
return None # fallback
Tip: before a flash sale or Black Friday, tell the vendor about the traffic you expect — most platforms will raise limits for the period.
Common mistakes with rate limits
- Not reading response headers — limits and remaining quota are in
X-RateLimit-*almost everywhere. - Retrying without backoff — an immediate repeat after a 429 reliably produces another 429.
- Ignoring limits until launch — discovering 429s in production under real load is expensive.
- No fallback — when the API is unavailable, recommendations should degrade gracefully to popular items or bestsellers rather than break the page.