Guide

AI crawlers: the search bots decide citations, the training bots do not

The AI crawlers OpenAI, Anthropic, Google and Perplexity document (OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, Googlebot, Google-Extended, PerplexityBot, Perplexity-User), what each one does, and which ones affect whether a site is cited.

Four vendors document their crawlers, and all four draw the same line: one set of agents decides whether a site appears in an assistant's search answers, another set collects content for model training, and the robots.txt settings are independent. A site that blocks "AI bots" wholesale usually means to opt out of training and accidentally opts out of being cited. The table is what each vendor says, read on October 8, 2026.

The crawlers, by vendor

VendorAgentWhat the vendor says it doesAffects citations?
OpenAIOAI-SearchBot"used to surface websites in search results in ChatGPT's search features"; sites that opt out "will not be shown in ChatGPT search answers"Yes
OpenAIChatGPT-Userfetches pages "when a user asks a question"; "not used for crawling the web in an automatic fashion"; "not used to determine whether content may appear in Search"Fetches at answer time
OpenAIGPTBotcrawls content "that may be used to train OpenAI's generative AI foundation models"; "each setting is independent of the others"No
AnthropicClaude-SearchBot"navigates the web to improve search result quality for users"; disabling it "may reduce your site's visibility and accuracy in user search results"Yes
AnthropicClaude-Userused when "individuals ask questions to Claude"; disabling it "prevents our system from retrieving your content in response to a user query"Fetches at answer time
AnthropicClaudeBotcollects "web content that could potentially contribute to their training"No
GoogleGooglebotcrawls for Google Search, "including Discover and all Google Search features"; AI Overviews and AI Mode need no additional requirements beyond Search eligibilityYes
GoogleGoogle-Extendeda control token for whether content is used "for training future generations of Gemini models" and for grounding; "does not impact a site's inclusion in Google Search nor is it used as a ranking signal"No for Search; grounding for Gemini
PerplexityPerplexityBot"designed to surface and link websites in search results on Perplexity"; "not used to crawl content for AI foundation models"Yes
PerplexityPerplexity-User"supports user actions within Perplexity"; "generally ignores robots.txt rules"; "not used for web crawling or to collect content for training"Fetches at answer time

Sources: developers.openai.com/docs/bots; support.claude.com (crawler article); developers.google.com (crawler list, "AI features and your website", the generative AI guide); docs.perplexity.ai (Perplexity crawlers). All read October 8, 2026; quotations as published.

The distinction that matters

The search crawlers (OAI-SearchBot, Claude-SearchBot, Googlebot, PerplexityBot) build the indexes the assistants search when a buyer asks a question. The user agents (ChatGPT-User, Claude-User, Perplexity-User) fetch a page at answer time because a user's question led there, and two of the three vendors say those fetches may not honor robots.txt because the user asked. The training crawlers (GPTBot, ClaudeBot, and the Google-Extended control) collect content for future models, and every vendor says blocking them does not affect search results. So:

  • Blocking GPTBot, ClaudeBot or Google-Extended opts a site out of training and leaves its citations alone.
  • Blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot removes a site from that assistant's search answers, which is the opposite of what most brands want.
  • Blocking Googlebot removes a site from Google Search, AI Overviews and AI Mode together.
  • A wildcard rule written for "AI bots" usually catches all of them at once.

What we check

Onboarding includes a read of the client's robots.txt against this table, because a blanket block is the one fix that costs nothing and changes everything. The weekly check then shows the result in the cited count: whether each assistant read a page on the domain for each question. A site that allows the search crawlers and is still not cited has a content problem, which is the rest of the program; a site that is blocking them has a one-line problem first.

Crawler analytics

Several tracking tools now report which AI bots hit which pages, how often, and the status code returned, from CDN or server logs. It is useful instrumentation and it answers a narrow question: did the crawler come. It does not answer whether the page was quoted, which only the answers do. Rightcited reports the answers; if a client has crawler logs we read them alongside, and the search-bot entries are the ones that matter.

Questions

Asked straight.

Should I block AI training crawlers?

That is a policy decision for your company, and it does not affect whether you are cited: every vendor says the training control is separate from search. If you block training, block only GPTBot, ClaudeBot and Google-Extended, and leave the search crawlers alone.

Does llms.txt help crawlers find my pages?

No vendor documents reading it as a ranking or citation signal, and Google says no new machine-readable files are needed for its AI features. We publish one because it is cheap and some assistants read it when asked about a company; it is not the lever.

How do I know which bots are reaching my site?

Your server or CDN logs list the user agents; the search bots in the table above are the ones to look for. A tracker's crawler analytics report is the same data presented for you. Either way, the answer that matters is whether the page was quoted, which the weekly check shows.

More guides

Read next.

See what the four assistants say about your brand this week.

Or start smaller: the AI Page Check runs one page through all four assistants for $2.99.

Get your free 360° AI Search Scan