What crawl budget is
Search bots cannot crawl the entire internet without limits — they allocate resources between
sites. Crawl budget is the number of pages a bot assigns to one specific site per period, usually a
day or a week.
The budget comes from two components:
- Crawl rate limit — how fast the bot can crawl the site without overloading the server.
- Crawl demand — how valuable the bot considers the pages, based on authority and how often
they change.
Why it matters in e-commerce
An online store with 200,000 SKUs generates millions of URLs through facets, sorting, pagination
and parameters. If the bot spends most of its budget on technical duplicates, new product and
category pages can wait weeks for indexing.
The situation:
- Real pages: 250,000
- Technical duplicates (facets + sorting): 1,800,000
- Crawl budget: 100,000 pages/day
- Result: new products wait 2–3 weeks for indexing
The main consumers of crawl budget
| Source | What it produces | The fix |
|---|---|---|
| Faceted search | Thousands of URLs like ?color=red&size=M | Canonical plus noindex, or a robots.txt disallow |
| Pagination | /page/1, /page/2… | Canonical to the first page, or paginated self-canonicals |
| Sorting parameters | ?sort=price_asc | Treat as duplicates: canonical to the unsorted URL |
| UTM and advertising parameters | ?utm_source=… | robots.txt disallow, or a canonical |
| Redirect chains | 301→302→200 | Shorten to a single hop |
| Internal search result pages | /search?q=… | Close with noindex |
Important: use the crawl stats and indexing reports in Google Search Console to monitor which
pages the bot actually crawls and how often.
How to spend the budget better
- Close off technical URLs through robots.txt or noindex — start with parameter pages.
- Place canonicals — on pages with similar content, state which URL is the canonical one.
- Optimise speed — bots crawl fast sites more actively; a TTFB above one second lowers the
crawl rate. - Keep the sitemap current — include only indexable, unique pages, with priorities set.
- Fix the link structure — pages with no internal links, known as orphan pages, are crawled
rarely.