AI search

How Claude picks what it cites

Only what the vendor documents, dated October 8, 2026, then what our weekly checks see Claude actually read, and what a business does about it.

Checked weeklyChatGPTClaudeGeminiPerplexity

What the vendor documents

Anthropic documents three agents and says its bots "respect 'do not crawl' signals by honoring industry standard directives in robots.txt":

  • Claude-SearchBot "navigates the web to improve search result quality for users"; disabling it "may reduce your site's visibility and accuracy in user search results".
  • Claude-User is the agent Claude uses when "individuals ask questions to Claude, it may access websites"; disabling it "prevents our system from retrieving your content in response to a user query, which may reduce your site's visibility for user-directed web search".
  • ClaudeBot collects "web content that could potentially contribute to their training"; a restriction "signals that the site's future materials should be excluded from our AI model training datasets".

Anthropic says opting out "requires modifying the robots.txt file" and that IP-based blocking "may not work correctly or persistently guarantee an opt-out". Claude's web search returns answers with citations to the pages it read; Anthropic does not document how those pages are ranked.

Source: support.claude.com, "Does Anthropic crawl data from the web, and how can site owners block the crawler?", read October 8, 2026.

What it means for your pages

Claude with search on behaves like a careful researcher with a budget. It runs a handful of searches per answer, reads what they return, and quotes with citations; when the budget runs out it says so and answers from what it has. For a business that means two things: the page that answers the question must be reachable by an ordinary search for the question, and the first pages Claude reaches set the answer. Our checks regularly see Claude answer a brand question from the two oldest listings about a company because its searches for the newer ones did not land before the limit.

From our checks

What we have seen Claude do.

Rows from scans and checks run in September and October 2026, with the companies unnamed where they have not agreed to be named.

CategoryWhat happened
HospitalityAsked about a resort group with four new properties, Claude ran nine searches, four of them for one new resort, hit its search limit, and answered from the group's two oldest villas on two review sites: "I wasn't able to retrieve specific review data" for the properties the buyer asked about.
ResortsAsked for the best honeymoon resorts in a region, Claude named three of the client's resorts fifth, sixth and tenth of ten, every line read from the client's own day-club blog; a second client page was quoted for a rival. Two of the five sources were rival resorts' own pages.
Software, pricingAsked which vendors price per trip, Claude named two rivals and four per-vehicle vendors from a competitor's pricing guide and three other sites, and never reached the client's pricing page, which ChatGPT had read twice the same morning.
Leather goodsAsked whether a brand was legit, Claude called it "this California-based company" off a third-party review (it is in Arizona), quoted a reviewer saying "top-grain" (the brand sells full-grain), ran five searches for where the goods are made, never reached the brand's site, and wrote "my searches were cut short".
ArchitectureAcross six answers Claude never opened the client's own site; every line it had about the firm came from its Houzz profile, and on one regional list that carried the firm with no city, Claude named eight other firms.
What we do

The work, for Claude.

The 50 questions run on Claude with web search every week, and the archive keeps the searches it ran as well as the pages it read, which is how we know when it stopped. The work for Claude is reachability: the page that answers the question has to come back for an ordinary search for that question, carry the fact in plain text, and exist on the sites Claude reaches first for your category. Where Claude answers from a directory instead of your site, the directory entry gets corrected and the site gets the text it was missing.

Questions

Asked straight.

Claude said it ran out of searches. Does that count?

Yes. A buyer who asks gets that answer, so we grade it as served and keep Claude's own sentence about the limit in the archive. We never re-run a question to get a better answer; the next weekly check is the re-run.

Does blocking ClaudeBot affect Claude's search answers?

Anthropic documents ClaudeBot as the training crawler and Claude-SearchBot and Claude-User as the search and fetch agents, with separate robots.txt controls. Blocking the search and user agents is what the vendor says may reduce visibility in search results.

Why did Claude read Houzz instead of our site?

Because the directory had text and the site had pictures. Claude quotes what it can read; a site that states its markets, projects and facts in words gives it something to quote, and a directory entry that is wrong gives it the wrong thing.

The other assistants

Read next.

See what the four assistants say about your brand this week.

Or start smaller: the AI Page Check runs one page through all four assistants for $2.99.

Get your free 360° AI Search Scan