How crawling works

A search bot — Googlebot, Bingbot or another crawler — starts from a set of known URLs, a kind of
seed list. Moving from page to page through links, it downloads the HTML, records status codes
(200, 301, 404, 500) and hands the content to the indexing system. That whole process is crawling.

Crawl rate is self-regulating: the bot watches how fast the server responds and lowers its request
frequency when the site is slow. Google does this automatically. Yandex, the dominant search engine
in Russia and several CIS markets, additionally allows the crawl rate to be set by hand in its
webmaster console.

What drives crawl efficiency in e-commerce

A large online store has potentially millions of URLs — products, filters, sort orders, pagination,
parameter combinations. Left unmanaged, the bot scatters its budget across low-value pages and
never gets round to new products or important categories.

The main factors:

Factor Effect on crawling
Server response time A poor TTFB slows the walk and shrinks the budget
Redirect chains Every redirect costs budget; three or more in a row may be ignored
Internal linking Pages with no inbound links are effectively invisible to the bot
XML sitemap Helps bots find new and updated URLs faster
Parameter URLs Thousands of ?sort=price&color=red create duplicates — close them with robots.txt or canonical

Managing the crawl zone

Not every page needs to be scanned. Keep these out of the crawl:

  • Filter and sort pages (?sort=, ?page=, ?color=)
  • Basket, account area, checkout
  • Technical endpoints (/api/, /admin/)
  • Duplicate versions of content (print views, AMP twins)

The instruments: robots.txt (a ban at the URL-pattern level), the noindex meta tag (allow the
scan, forbid the index entry) and canonical (a signal about the preferred URL when duplicates
exist).

Common problems

The bot falls into a trap. Dynamically generated pages — on-site search results, infinite
pagination — can spawn hundreds of thousands of URLs. Close them in robots.txt.

JavaScript content is invisible. Recommendation widgets, prices and descriptions that load
through JavaScript after page initialisation reach the bot late, or never. Business-critical
content belongs in the initial HTML response.

New product pages are not crawled in time. The fix: add the <url> entry to the sitemap with a
<lastmod> value when a product is published — that is a priority signal for the bot.

Important: crawling is not indexation. A crawled page may still fail to enter the index if its
content is judged duplicate or low quality. Check the status of specific URLs in Google Search
Console, and in Yandex Webmaster if the Russian-language market matters to you.