What rate limiting is for

Rate limiting is a policy that caps how many requests a client may make in a given period. The goals:

  • Protection from overload — preventing service degradation when traffic spikes.
  • Fair resource sharing — one client cannot consume all the capacity.
  • Protection from abuse — denial-of-service attacks, scraping, credential stuffing.

When the limit is exceeded, the API answers:

HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 1000
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1716300000

Rate limiting algorithms

Algorithm Principle Characteristic
Fixed window A counter over a fixed period (a minute, an hour) Bursts at the window boundary
Sliding window A rolling window with no fixed boundaries More precise, no boundary spikes
Token bucket A bucket refilled at a set rate Tolerates bursts up to the bucket size
Leaky bucket A queue draining at a fixed rate Strictly even processing

Rate limiting in e-commerce integrations

An online store running a personalization platform calls the API on every page view: the product page requests recommendations, the listing requests ranking, the home page requests banners. At 100K unique visitors a day that is roughly 3 to 5 million API calls in 24 hours, with peaks of several thousand requests per second.

Practical rules for integrations:

  1. Cache client-side. Recommendations for a given user identifier rarely change within one session — a cache with a 5 to 15 minute TTL cuts request volume threefold to fivefold.
  2. Use batch requests. Where the API supports it, fetch recommendations for several widgets in one call.
  3. Circuit breaker. After a run of 429 errors, switch automatically to a fallback — popular products from a local cache.
  4. Exponential backoff. Never retry immediately after a 429.
import time

def api_request_with_retry(url, max_retries=3):
    for attempt in range(max_retries):
        response = requests.get(url)
        if response.status_code == 429:
            wait = 2 ** attempt  # 1, 2, 4 seconds
            time.sleep(wait)
            continue
        return response
    return None  # fallback

Tip: before a flash sale or Black Friday, tell the vendor about the traffic you expect — most platforms will raise limits for the period.

Common mistakes with rate limits

  • Not reading response headers — limits and remaining quota are in X-RateLimit-* almost everywhere.
  • Retrying without backoff — an immediate repeat after a 429 reliably produces another 429.
  • Ignoring limits until launch — discovering 429s in production under real load is expensive.
  • No fallback — when the API is unavailable, recommendations should degrade gracefully to popular items or bestsellers rather than break the page.