What crawl budget is

Search bots cannot crawl the entire internet without limits — they allocate resources between
sites. Crawl budget is the number of pages a bot assigns to one specific site per period, usually a
day or a week.

The budget comes from two components:

  • Crawl rate limit — how fast the bot can crawl the site without overloading the server.
  • Crawl demand — how valuable the bot considers the pages, based on authority and how often
    they change.

Why it matters in e-commerce

An online store with 200,000 SKUs generates millions of URLs through facets, sorting, pagination
and parameters. If the bot spends most of its budget on technical duplicates, new product and
category pages can wait weeks for indexing.

The situation:
- Real pages: 250,000
- Technical duplicates (facets + sorting): 1,800,000
- Crawl budget: 100,000 pages/day
- Result: new products wait 2–3 weeks for indexing

The main consumers of crawl budget

Source What it produces The fix
Faceted search Thousands of URLs like ?color=red&size=M Canonical plus noindex, or a robots.txt disallow
Pagination /page/1, /page/2… Canonical to the first page, or paginated self-canonicals
Sorting parameters ?sort=price_asc Treat as duplicates: canonical to the unsorted URL
UTM and advertising parameters ?utm_source=… robots.txt disallow, or a canonical
Redirect chains 301→302→200 Shorten to a single hop
Internal search result pages /search?q=… Close with noindex

Important: use the crawl stats and indexing reports in Google Search Console to monitor which
pages the bot actually crawls and how often.

How to spend the budget better

  1. Close off technical URLs through robots.txt or noindex — start with parameter pages.
  2. Place canonicals — on pages with similar content, state which URL is the canonical one.
  3. Optimise speed — bots crawl fast sites more actively; a TTFB above one second lowers the
    crawl rate.
  4. Keep the sitemap current — include only indexable, unique pages, with priorities set.
  5. Fix the link structure — pages with no internal links, known as orphan pages, are crawled
    rarely.