What robots.txt is and how it works

The file /robots.txt is the first thing most search crawlers request before fetching a site. It holds instructions on which sections may be crawled and which may not. The standard is described in RFC 9309, the Robots Exclusion Protocol.

The file is made of blocks, each starting with a User-agent directive (who) followed by Allow and Disallow lines (what):

User-agent: *
Disallow: /cart/
Disallow: /checkout/
Disallow: /account/
Disallow: /search?
Allow: /search/

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap_index.xml

* means all bots. A specific User-agent block overrides the rules for that particular bot.

The key limitation: crawling is not indexation

robots.txt governs crawling — whether a bot visits a page. It does not govern indexation directly.

What you need The tool
Do not crawl the page Disallow in robots.txt
Do not index the page <meta name="noindex"> or X-Robots-Tag: noindex
Neither crawl nor index The noindex tag (robots.txt will not help — a bot that is blocked from crawling never sees the meta tag)

A page closed with Disallow can still end up in the index through external links: Googlebot learns of it, never visits, and indexes it as a URL with no content.

What to block in e-commerce

Always block:
– /cart/, /checkout/, /account/, /login/ — utility pages with no SEO value
– /search?q= — internal search results generate thousands of duplicate URLs
– Sort and filter parameters where they create duplicates (?sort=price, ?color=red&size=M)
– UTM parameters (?utm_source=, ?utm_medium=)

Do not block:
– Category pages with SEO value
– Product pages
– Static landing pages, the blog, case studies

Important: blocking in robots.txt does not remove pages that are already indexed. To take existing URLs out, use the removal tool in Google Search Console or a noindex tag.

robots.txt for AI crawlers

Between 2023 and 2025 bots from OpenAI, Anthropic, Perplexity, Apple and other AI companies appeared. They respect robots.txt, but by default they are allowed to crawl. If you want to close your content off from LLM training:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: PerplexityBot
Disallow: /

The opposite strategy, for AEO and GEO: explicitly open your expert content — documentation, case studies, glossaries — to AI crawlers so it gets cited in answers from ChatGPT, Perplexity and Claude.

Common mistakes

  • Blocking CSS and JS files. If Googlebot cannot load a page’s styles and scripts, it cannot see how the page renders and may rank it lower.
  • Blocking the whole site by accident. Disallow: / stops all crawling — a common error when setting up a staging environment. Check robots.txt separately on production and on dev.
  • Treating robots.txt as protection. The file is public and readable by anyone at its direct URL. Confidential content is protected by authentication, not by robots.txt.