If an AI system's crawler cannot reach a page, that page cannot be retrieved or cited. Access is the cheapest problem to fix and one of the most common.
Know the crawlers
Each company runs different crawlers for different jobs. As of October 2026, examples include:
- Googlebot: Google Search, including AI Overviews and AI Mode.
- Google-Extended: a control token for Gemini model use, not a separate crawler, and not used for Search ranking.
- Bingbot: Bing, and so Copilot.
- OAI-SearchBot: ChatGPT search. GPTBot is the training crawler.
- PerplexityBot: Perplexity.
- ClaudeBot: Anthropic.
Check each company's documentation for the current names before you write rules.
Check four places
- robots.txt, for rules blocking any of these agents.
- Your CDN or firewall, where bot protection often blocks AI crawlers without anyone deciding to.
- Server logs, to confirm the crawlers are actually visiting and getting 200 responses.
- Login walls and cookie gates that hide content.
Decide deliberately
You may have good reasons to block training crawlers. That is a business decision. Blocking search crawlers by accident is not. Write the policy down and review it twice a year.
About llms.txt
llms.txt is a proposed file listing your key content for language models. It is cheap to add, but no major AI search platform has confirmed it uses it for ranking or citation. Do not prioritise it over crawl access.