What AI crawlers are

AI crawlers are robots run by the companies that build language models and AI assistants. Technically they do the same as search bots during crawling — they download pages. The difference is what happens to the content next: it may go into a training dataset, into an AI search index or straight into an answer for the person who asked a question. That is why a single allow-or-block-AI rule is not enough; the decision is made for each type separately.

Three types of AI bots

Purpose OpenAI Anthropic Perplexity Google
Model training GPTBot ClaudeBot none declared Google-Extended (robots.txt token)
AI search index OAI-SearchBot Claude-SearchBot PerplexityBot —
User request ChatGPT-User Claude-User Perplexity-User Google-Agent

The names follow the companies’ documentation as of September 2026. Google-Extended differs from the rest: it is a token, not a separate bot. Crawling is done by Google’s regular crawlers, and the token decides whether the content may be used to train Gemini and for grounding. User-triggered fetchers are better classed as AI agent traffic: each visit is driven by a specific person’s request.

How much they take and how much they send back

Cloudflare compares how often a company’s crawler requests pages with how many referrals its product sends back. For July 2025 the ratios were:

Company Crawler requests per referral
Anthropic 38,065
OpenAI 1,091
Perplexity 194
Google 5.4

Over the 12 months to July 2025, 80% of AI bot crawling was for training, 18% for search and 2% for user-requested actions.

A sample robots.txt for a store

An option for a store that wants links in AI answers (GEO) but does not want to hand its content over for model training: block training, allow AI search.

# Model training - blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

# AI search - allowed, except service sections
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /cart/
Disallow: /checkout/
Allow: /

# YandexGPT and Search with Alice answers - a separate decision
User-agent: YandexAdditional
Allow: /

Important: blocking Google-Extended also opts you out of grounding in Gemini, not just training. Make that decision deliberately.

Checklist

  • Find AI bots in your server logs and check their IPs against the companies’ published lists: the User-Agent name is easy to fake, and not every bot carries Web Bot Auth signatures yet.
  • Decide by type: training, search and user requests are different business questions.
  • Remember that user-triggered fetchers may not follow robots.txt: if you need to restrict them, do it at the CDN or WAF level.
  • After editing robots.txt, check its syntax and give bots time: OpenAI cites about 24 hours.