The one thing to understand
Every major AI company runs at least two different crawlers. One gathers training data. A different one decides whether you can be cited in that assistant's answers. They are separate user-agents, and they are separate business decisions. Confusing them is the single most expensive robots.txt mistake being made right now.
Three jobs, three kinds of crawler
Before touching any config, it helps to see the shape of the system. AI crawlers fall into three lanes, and only one of them affects whether an assistant can cite you.
The full crawler reference
| User-agent | Company | Purpose | Block = ? |
|---|---|---|---|
| Googlebot | Search + AI Overviews + AI Mode | Invisible in Google entirely | |
| Google-Extended | Gemini grounding, model training | No effect on AI Overviews | |
| OAI-SearchBot | OpenAI | ChatGPT search inclusion | Gone from ChatGPT search |
| GPTBot | OpenAI | Model training | Opts out of training only |
| ChatGPT-User | OpenAI | Live fetch on user request | Users cannot open your links |
| Claude-SearchBot | Anthropic | Search result quality | Gone from Claude answers |
| Claude-User | Anthropic | Live fetch on user request | Users cannot open your links |
| ClaudeBot | Anthropic | Model training | Opts out of training only |
| PerplexityBot | Perplexity | Search index | Gone from Perplexity |
| Applebot-Extended | Apple | Apple model training | Training opt-out only |
| meta-externalagent | Meta | Meta AI training and products | Training opt-out |
The Google-Extended trap
A lot of Nepal sites added Google-Extended: Disallow believing it removes them from AI Overviews. It does not. Google-Extended governs Gemini grounding and training. AI Overviews and AI Mode run on Googlebot, because Google states AI is built into Search itself. If you genuinely want out of AI Overviews, use the Search Console generative AI toggle or page-level nosnippet, not Google-Extended.
Copy-paste: allow AI citation (recommended for most businesses)
For a service business, being quoted in an AI answer is free distribution to someone actively researching what you sell. This config allows everything, and is what I run on this site.
User-agent: *
Allow: /
# Google (AI Overviews + AI Mode ride on Googlebot)
User-agent: Googlebot
Allow: /
# OpenAI: search inclusion, then training, then user fetch
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Allow: /
User-agent: ChatGPT-User
Allow: /
# Anthropic: three separate agents
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: ClaudeBot
Allow: /
# Perplexity
User-agent: PerplexityBot
Allow: /
Sitemap: https://yourdomain.com/sitemap.xml
Copy-paste: stay citable, opt out of training
This is the balanced position: your content can still be found and cited, but is not collected for model training. Useful if you publish original research or proprietary analysis.
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Opt out of model training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
Rate-limiting instead of blocking
If the problem is server load rather than principle, Anthropic supports the non-standard Crawl-delay extension:
Crawl-delay: 1
Anthropic also notes that blocking by IP address is unreliable, because it can prevent the crawler reading your robots.txt in the first place, and they do not publish fixed IP ranges since they use provider IPs.
How to actually measure AI traffic
Allowing the crawlers is pointless if you never check whether it produced anything. The good news is ChatGPT tags its referrals automatically.
In Google Analytics 4
ChatGPT appends utm_source=chatgpt.com to outgoing links, so no setup is needed.
Go to Reports → Acquisition → Traffic acquisition, then switch the dimension to Session source / medium and look for chatgpt.com. Perplexity and Claude typically arrive as ordinary referrals from their own domains.
In Search Console
Google added generative AI performance reporting in June 2026, showing impressions and which pages appear in AI responses, by country.
It rolled out in stages by country, and currently reports impressions rather than clicks, so treat it as a visibility signal rather than a traffic number.
Four mistakes worth avoiding
Blocking the search bot while allowing the training bot
The exact inverse of what almost everyone intends. You still contribute to training and lose all citation visibility.
Assuming Google-Extended controls AI Overviews
It does not. It governs Gemini and training. AI Overviews follow Googlebot and the Search Console toggle.
Using only the legacy anthropic-ai token
Anthropic's current agents are ClaudeBot, Claude-SearchBot and Claude-User. A robots.txt listing only the old token misses the two that matter.
Blocking by IP or firewall rule
Cloudflare bot rules and IP blocks frequently catch legitimate crawlers and, as Anthropic notes, can stop a crawler reading robots.txt at all. Control access in robots.txt, not at the firewall.
One closing caveat: robots.txt governs whether a crawler may fetch you, not whether an assistant will choose to cite you. That second question is about whether the content deserves citing, which is a content problem rather than a config problem. I dug into what Google says actually drives that in AI SEO myths: what Google actually says.
Written by someone who does this for a living
I run SEO, ads and the development behind them
I work with US-based companies on technical SEO and Next.js builds. On one, a 90-day campaign grew organic clicks 271% and AI-assistant referrals 325%.
Frequently Asked Questions
What is the difference between GPTBot and OAI-SearchBot?
GPTBot collects content that may train OpenAI models. OAI-SearchBot makes your site eligible to appear and be cited in ChatGPT search. Blocking GPTBot opts you out of training while keeping search visibility. Blocking OAI-SearchBot removes you from ChatGPT search results.
Does blocking Google-Extended remove me from AI Overviews?
No. Google-Extended controls Gemini grounding and training, not AI Overviews or AI Mode. Those run on Googlebot because AI is built into Search. Use the Search Console generative AI toggle, or nosnippet and max-snippet at page level.
What are Anthropic's three crawlers?
ClaudeBot collects content that may contribute to training. Claude-SearchBot improves search result quality and is the one that matters for citation. Claude-User fetches a page when a user asks about it. All three respect robots.txt and each has its own user-agent block.
How do I track ChatGPT traffic in Google Analytics?
ChatGPT automatically appends utm_source=chatgpt.com to referral URLs, so it appears in GA4 with no setup. Open Reports, then Acquisition, then Traffic acquisition, and switch the dimension to Session source / medium to find chatgpt.com.
Should a small business block AI crawlers?
Usually no. Being cited is free distribution to someone actively researching your service. Blocking mainly suits publishers who monetise the pageview itself. A middle path is allowing the search crawlers while blocking the training crawlers.
Sources
- Overview of OpenAI Crawlers
- Anthropic: does Anthropic crawl data from the web, and how can site owners block the crawler
- AI Features and Your Website, Google Search Central
- OpenAI: Searching the web with ChatGPT
Crawler names verified against vendor documentation on 23 August 2026. These change, so re-check before relying on an old config.