Technical SEO

AI Crawlers and robots.txt: The Complete 2026 Guide

Most sites that tried to "block AI" in the last two years blocked the wrong bot. They kept the training crawler and removed the search crawler, which is precisely backwards: they still fed the models and simply deleted themselves from the answers. Here is every crawler that matters, what each one actually controls, and a config you can paste today.

The one thing to understand

Every major AI company runs at least two different crawlers. One gathers training data. A different one decides whether you can be cited in that assistant's answers. They are separate user-agents, and they are separate business decisions. Confusing them is the single most expensive robots.txt mistake being made right now.

Three jobs, three kinds of crawler

Before touching any config, it helps to see the shape of the system. AI crawlers fall into three lanes, and only one of them affects whether an assistant can cite you.

The three types of AI crawler Training crawlers such as GPTBot, ClaudeBot and Google-Extended feed model training and do not affect citation. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot decide whether a site can be cited in AI answers. User-fetch crawlers such as ChatGPT-User and Claude-User retrieve a page live when a person asks about it. 1. Training crawlers Collect content that may train future models. Blocking these does NOT hide you from AI answers. GPTBot ClaudeBot Google-Extended Applebot-Extended CCBot 2. Search crawlers: these decide citation Build the index the assistant answers from. Block one of these and you vanish from that assistant. OAI-SearchBot Claude-SearchBot PerplexityBot Googlebot bingbot 3. User-fetch crawlers Fetch one page live because a person asked about it. Blocking these breaks "read this link for me". ChatGPT-User Claude-User Perplexity-User
Only the green lane controls whether an AI assistant can cite you. Blocking the amber lane opts you out of training without hurting visibility.

The full crawler reference

User-agentCompanyPurposeBlock = ?
GooglebotGoogleSearch + AI Overviews + AI ModeInvisible in Google entirely
Google-ExtendedGoogleGemini grounding, model trainingNo effect on AI Overviews
OAI-SearchBotOpenAIChatGPT search inclusionGone from ChatGPT search
GPTBotOpenAIModel trainingOpts out of training only
ChatGPT-UserOpenAILive fetch on user requestUsers cannot open your links
Claude-SearchBotAnthropicSearch result qualityGone from Claude answers
Claude-UserAnthropicLive fetch on user requestUsers cannot open your links
ClaudeBotAnthropicModel trainingOpts out of training only
PerplexityBotPerplexitySearch indexGone from Perplexity
Applebot-ExtendedAppleApple model trainingTraining opt-out only
meta-externalagentMetaMeta AI training and productsTraining opt-out

The Google-Extended trap

A lot of Nepal sites added Google-Extended: Disallow believing it removes them from AI Overviews. It does not. Google-Extended governs Gemini grounding and training. AI Overviews and AI Mode run on Googlebot, because Google states AI is built into Search itself. If you genuinely want out of AI Overviews, use the Search Console generative AI toggle or page-level nosnippet, not Google-Extended.

Copy-paste: allow AI citation (recommended for most businesses)

For a service business, being quoted in an AI answer is free distribution to someone actively researching what you sell. This config allows everything, and is what I run on this site.

# Allow AI assistants to find and cite this site
User-agent: *
Allow: /

# Google (AI Overviews + AI Mode ride on Googlebot)
User-agent: Googlebot
Allow: /

# OpenAI: search inclusion, then training, then user fetch
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Allow: /
User-agent: ChatGPT-User
Allow: /

# Anthropic: three separate agents
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: ClaudeBot
Allow: /

# Perplexity
User-agent: PerplexityBot
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

Copy-paste: stay citable, opt out of training

This is the balanced position: your content can still be found and cited, but is not collected for model training. Useful if you publish original research or proprietary analysis.

# Keep search visibility
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /

# Opt out of model training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /

Rate-limiting instead of blocking

If the problem is server load rather than principle, Anthropic supports the non-standard Crawl-delay extension:

User-agent: ClaudeBot
Crawl-delay: 1

Anthropic also notes that blocking by IP address is unreliable, because it can prevent the crawler reading your robots.txt in the first place, and they do not publish fixed IP ranges since they use provider IPs.

How to actually measure AI traffic

Allowing the crawlers is pointless if you never check whether it produced anything. The good news is ChatGPT tags its referrals automatically.

In Google Analytics 4

ChatGPT appends utm_source=chatgpt.com to outgoing links, so no setup is needed.

Go to Reports → Acquisition → Traffic acquisition, then switch the dimension to Session source / medium and look for chatgpt.com. Perplexity and Claude typically arrive as ordinary referrals from their own domains.

In Search Console

Google added generative AI performance reporting in June 2026, showing impressions and which pages appear in AI responses, by country.

It rolled out in stages by country, and currently reports impressions rather than clicks, so treat it as a visibility signal rather than a traffic number.

Four mistakes worth avoiding

Blocking the search bot while allowing the training bot

The exact inverse of what almost everyone intends. You still contribute to training and lose all citation visibility.

Assuming Google-Extended controls AI Overviews

It does not. It governs Gemini and training. AI Overviews follow Googlebot and the Search Console toggle.

Using only the legacy anthropic-ai token

Anthropic's current agents are ClaudeBot, Claude-SearchBot and Claude-User. A robots.txt listing only the old token misses the two that matter.

Blocking by IP or firewall rule

Cloudflare bot rules and IP blocks frequently catch legitimate crawlers and, as Anthropic notes, can stop a crawler reading robots.txt at all. Control access in robots.txt, not at the firewall.

One closing caveat: robots.txt governs whether a crawler may fetch you, not whether an assistant will choose to cite you. That second question is about whether the content deserves citing, which is a content problem rather than a config problem. I dug into what Google says actually drives that in AI SEO myths: what Google actually says.

Rahul Ranjan, SEO specialist and full stack developer

Written by someone who does this for a living

I run SEO, ads and the development behind them

I work with US-based companies on technical SEO and Next.js builds. On one, a 90-day campaign grew organic clicks 271% and AI-assistant referrals 325%.

Frequently Asked Questions

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects content that may train OpenAI models. OAI-SearchBot makes your site eligible to appear and be cited in ChatGPT search. Blocking GPTBot opts you out of training while keeping search visibility. Blocking OAI-SearchBot removes you from ChatGPT search results.

Does blocking Google-Extended remove me from AI Overviews?

No. Google-Extended controls Gemini grounding and training, not AI Overviews or AI Mode. Those run on Googlebot because AI is built into Search. Use the Search Console generative AI toggle, or nosnippet and max-snippet at page level.

What are Anthropic's three crawlers?

ClaudeBot collects content that may contribute to training. Claude-SearchBot improves search result quality and is the one that matters for citation. Claude-User fetches a page when a user asks about it. All three respect robots.txt and each has its own user-agent block.

How do I track ChatGPT traffic in Google Analytics?

ChatGPT automatically appends utm_source=chatgpt.com to referral URLs, so it appears in GA4 with no setup. Open Reports, then Acquisition, then Traffic acquisition, and switch the dimension to Session source / medium to find chatgpt.com.

Should a small business block AI crawlers?

Usually no. Being cited is free distribution to someone actively researching your service. Blocking mainly suits publishers who monetise the pageview itself. A middle path is allowing the search crawlers while blocking the training crawlers.

Sources

Crawler names verified against vendor documentation on 23 August 2026. These change, so re-check before relying on an old config.

Rahul Ranjan, technical SEO specialist in Nepal

Rahul Ranjan

SEO & Digital Growth Specialist · Full Stack Developer · Biratnagar, Nepal

I manage SEO, analytics, and lead generation, and I implement fixes in code rather than handing over a report. Crawler configuration is one of the few SEO areas where a single wrong line has an immediate, total effect.

Chat on WhatsApp