What AI crawlers are
AI crawlers are robots run by the companies that build language models and AI assistants. Technically they do the same as search bots during crawling — they download pages. The difference is what happens to the content next: it may go into a training dataset, into an AI search index or straight into an answer for the person who asked a question. That is why a single allow-or-block-AI rule is not enough; the decision is made for each type separately.
Three types of AI bots
| Purpose | OpenAI | Anthropic | Perplexity | |
|---|---|---|---|---|
| Model training | GPTBot | ClaudeBot | none declared | Google-Extended (robots.txt token) |
| AI search index | OAI-SearchBot | Claude-SearchBot | PerplexityBot | — |
| User request | ChatGPT-User | Claude-User | Perplexity-User | Google-Agent |
The names follow the companies’ documentation as of September 2026. Google-Extended differs from the rest: it is a token, not a separate bot. Crawling is done by Google’s regular crawlers, and the token decides whether the content may be used to train Gemini and for grounding. User-triggered fetchers are better classed as AI agent traffic: each visit is driven by a specific person’s request.
How much they take and how much they send back
Cloudflare compares how often a company’s crawler requests pages with how many referrals its product sends back. For July 2025 the ratios were:
| Company | Crawler requests per referral |
|---|---|
| Anthropic | 38,065 |
| OpenAI | 1,091 |
| Perplexity | 194 |
| 5.4 |
Over the 12 months to July 2025, 80% of AI bot crawling was for training, 18% for search and 2% for user-requested actions.
A sample robots.txt for a store
An option for a store that wants links in AI answers (GEO) but does not want to hand its content over for model training: block training, allow AI search.
# Model training - blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
# AI search - allowed, except service sections
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /cart/
Disallow: /checkout/
Allow: /
# YandexGPT and Search with Alice answers - a separate decision
User-agent: YandexAdditional
Allow: /
Important: blocking Google-Extended also opts you out of grounding in Gemini, not just training. Make that decision deliberately.
Checklist
- Find AI bots in your server logs and check their IPs against the companies’ published lists: the User-Agent name is easy to fake, and not every bot carries Web Bot Auth signatures yet.
- Decide by type: training, search and user requests are different business questions.
- Remember that user-triggered fetchers may not follow robots.txt: if you need to restrict them, do it at the CDN or WAF level.
- After editing robots.txt, check its syntax and give bots time: OpenAI cites about 24 hours.