Undercover AI Agents · the register, applied to your own file

Your site has an AI policy.
You probably didn’t write it.

Hosting providers, CDNs, security plugins and SEO plugins all ship lists that block AI crawlers, and several apply them without asking. Put your domain in and see what yours actually says, what each part of it costs you, and whether a person chose any of it.

Reads one file, /robots.txt, over https. The domain you type is not logged, stored or counted — there is no analytics on this page.

Three decisions, not one

Nearly every tool treats “AI crawlers” as a single switch. They are not. The operators run separate agents for separate jobs, and refusing each one costs you something different.

Training

Collects pages that may train a future model. Blocking costs you no traffic at all. This is the one most people mean when they say they want to block AI.

GPTBot · ClaudeBot · CCBot · Bytespider

Search

Builds the pool an assistant cites from. Blocking means you stop being cited when someone asks a question your page answers.

OAI-SearchBot · Claude-SearchBot · PerplexityBot

Errand

Sent because a specific person asked about your specific page. Blocking tells that reader your page cannot be read. Almost nobody chooses this on purpose.

ChatGPT-User · Claude-User · Perplexity-User

Two names you will see on block lists are not crawlers at all. Google-Extended and Applebot-Extended are robots.txt opt-out tokens for training permission; nothing ever sends them as a user-agent. Blocking them is meaningful, but it is not blocking a visitor.