If a bot can't read your page, it doesn't index it or cite it. With Google you notice it in Search Console. With ChatGPT or Perplexity, you just stop appearing as a source and nobody tells you.

Almost always it's one of these four things.

1. The robots.txt

It's the first thing any serious bot reads. A Disallow: / in the User-agent: * group locks out every bot that doesn't have its own group. It happens a lot after a migration, when the staging robots.txt gets pushed to production.

One detail few people know: if the robots.txt returns a 5xx error, Google stops crawling the site until it responds again. If it returns a 404, on the other hand, Google understands it can crawl everything.

2. The firewall or CDN

Cloudflare, Akamai, DataDome and the like block bots by user-agent, by IP or by behavior. Cloudflare, for example, ships an anti-bot mode and an option to block AI crawlers that some sites have turned on without knowing it.

The tool compares what a normal browser receives with what each bot receives. If the browser sees the page and the bot gets a 403 or a challenge, something is treating them differently.

3. The noindex

The bot gets in, but you tell it not to index. It can come in a <meta name="robots"> tag or in the X-Robots-Tag header, and it can target a specific bot (googlebot: noindex).

4. Different content for bots

Some sites serve bots an empty page, an unrendered version or a cookie notice that covers everything. If a bot gets far less content than the browser, the tool flags it so you can take a look.

AI search engines: which bots to let in

OpenAI, Anthropic and Perplexity split their bots by function:

  • Search (OAI-SearchBot, Claude-SearchBot, PerplexityBot): index your site so it can be cited in their answers.
  • User (ChatGPT-User, Claude-User, Perplexity-User): visit when someone asks them to open your page.
  • Training (GPTBot, ClaudeBot, CCBot…): collect content to train models.

If you want to appear in ChatGPT, Claude or Perplexity, let in at least the search and user bots. Blocking the training bots is a business decision and doesn't affect whether you get cited.

Google is different: Google-Extended controls whether your content is used to train Gemini, but AI Overviews and AI Mode use the same Googlebot as the search engine. If you block Googlebot, you disappear from everything.

Limits of the tool

The requests leave from Cloudflare with each bot's user-agent. Sites that verify bots by IP may block you here and let the real bot through, and some even serve it a different robots.txt. So an HTTP block tells you where to look, but confirm it in your logs or with URL Inspection in Search Console.