Should you let AI crawlers in?
Blocking AI crawlers is a reasonable decision made badly almost everywhere, because the tokens people block are rarely the ones doing what they object to.
The distinction that decides everything
Every major provider runs separate agents for separate purposes, documented and individually controllable.
- OpenAI runs
GPTBotfor training,OAI-SearchBot“used to surface websites in search results in ChatGPT’s search features”,ChatGPT-Userfor user actions, andOAI-AdsBot(OpenAI). BlockingGPTBotdoes not remove you from ChatGPT search.OAI-SearchBotis the citation path. - Anthropic runs
ClaudeBotfor training,Claude-Userwhen a person asks Claude something, andClaude-SearchBotto “improve search result quality” (Anthropic). - Perplexity runs
PerplexityBot, which it says “is not used to crawl content for AI foundation models”, andPerplexity-User. Note this, from Perplexity’s own documentation: because a user requested the fetch,Perplexity-User“generally ignores robots.txt rules” (Perplexity). - Google uses
Googlebotfor Search, and AI Overviews and AI Mode are part of Search.Google-Extendedgoverns whether your content trains or grounds Gemini models. It is not a crawler, and it does not affect Search inclusion or ranking (Google). - Apple follows the same pattern: “Applebot-Extended does not crawl webpages. Webpages that disallow Applebot-Extended can still be included in search results” (Apple).
So the common configuration, block GPTBot and Google-Extended, opts you out
of training while leaving AI answer visibility untouched. That may be exactly
what you want. It is worth knowing it is what you chose.
To actually leave Google’s AI answers, the control is the Search Console Search generative AI setting, not a robots.txt line, as covered in the plumbing piece.
The economics people are reacting to
The objection is usually that AI crawlers take content and send nothing back. Cloudflare quantified it: for a window in August 2025, across all industries, it measured roughly 50,000 HTML requests per referral for Anthropic, 887:1 for OpenAI and 118:1 for Perplexity (Cloudflare). In that same period, training accounted for “nearly 80% of the crawling from AI bots”.
Two caveats Cloudflare itself raises. Traffic referred from Claude’s native app carries no referrer header, which inflates Anthropic’s ratio. And these are rolling-window metrics that have moved by orders of magnitude between windows, so any figure without its window attached is not a statistic.
What the web decided
The Data Provenance Initiative audited 14,000 domains and found restrictions climbing fast: among the most critical sources in the C4 corpus, “20-33% of all tokens are restricted, as compared to <3% one year prior”, with OpenAI’s crawlers blocked for 25.9% of tokens in that head distribution (Longpre et al., arXiv:2407.14933). Their data runs to April 2024, so read it as the start of a trend.
They also noted the asymmetry that matters commercially: some providers register training crawlers for opt-out while inference-time fetchers go unregistered.
A decision framework
Ask what you are protecting. If it is your content being used to train a competitor’s model, block the training tokens and keep the search tokens. If it is bandwidth, rate-limit rather than block. If your commercial goal is being recommended, be careful: a brand absent from both the index and the training data is one the assistant has neither retrieved nor remembered, and the answer will be built from whatever third parties say about you instead.