AI crawler checker: is your site blocking ChatGPT, Claude or Perplexity?
An AI engine can only name and cite a page its crawler can read. Enter your domain to see, for each major AI crawler, whether your robots.txt allows it and whether your server or CDN actually lets it in. Free, no signup, results in a few seconds.
| Crawler | What it does | robots.txt | Live request |
|---|
Crawlable is step one. Get a free visibility audit to see whether ChatGPT, Gemini, Perplexity and Google AI Overviews actually name you.
What the two columns mean
robots.txt is the instruction your site gives crawlers. The checker applies the rules crawlers follow: the group that names a crawler beats the catch-all User-agent: * group, and within a group the longest matching rule wins. "Blocked" means the crawler is told not to read your homepage; "allowed, some paths restricted" means the homepage is open but other paths are disallowed.
We read the public robots.txt, the one any visitor gets. A few large sites serve crawlers a different file from the public one, and no outside tool can see that version; if you run such a setup, your own server logs are the authority.
Live request is what happens when a request identifying as that crawler reaches your homepage. Googlebot and Bingbot are not tested this way, because sites rightly refuse anything claiming to be them that does not come from Google or Microsoft. If a normal browser gets the page and the crawler gets a 403 or 503, a firewall or CDN rule is refusing it, whatever robots.txt says. Our request does not come from the crawler's own IP addresses, so treat this column as a strong hint rather than proof.
Training, search and user crawlers are different choices
Most AI companies run more than one crawler. Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot) build the indexes AI answers cite, so blocking them removes you from those answers. User crawlers (ChatGPT-User, Claude-User, Perplexity-User) fetch a page when a person asks the assistant to read it. Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) collect content for future models. Many sites block training and allow search; that is a legitimate choice, as long as it is a choice and not an accident.
How to unblock a crawler
In robots.txt, give the crawler its own group. A crawler obeys only the most specific group that names it, so this overrides a blanket Disallow under User-agent: *:
User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: /
If robots.txt already allows the crawler but the live request is blocked, the rule lives in your CDN or firewall. On Cloudflare, look at the AI crawler and bot settings for the domain; other CDNs and security plugins have equivalents. For background, see how Cloudflare's AI crawler blocking works.
Crawlable is not the same as cited
Once the crawlers can read your site, the question is whether AI engines name you when buyers ask about your category. The free audit asks Google AI Overviews, ChatGPT, Gemini and Perplexity and shows where a competitor is named instead of you.
Get your free visibility auditFrequently asked questions
How do I check if my site blocks AI crawlers?
Enter your domain above. The checker reads your robots.txt and applies the same matching rules crawlers use, then sends a request to your homepage as each AI crawler and compares the response with a normal browser request. A blocked result in either column means that crawler cannot read your homepage.
Does blocking GPTBot remove my site from ChatGPT?
Not on its own. GPTBot collects content for training. ChatGPT's search answers are built by OAI-SearchBot, and pages a user asks ChatGPT to open are fetched by ChatGPT-User. To appear in ChatGPT search results, OAI-SearchBot must be allowed.
Does Google-Extended control AI Overviews?
No. Google-Extended controls whether your content is used for Gemini apps and Vertex AI. AI Overviews and AI Mode are part of Google Search and use Googlebot, so blocking Googlebot is the only robots.txt setting that removes you from them, and it removes you from Search as well.
Why does robots.txt say allowed but the live request is blocked?
Something in front of your site, usually a CDN or firewall rule, is refusing the crawler before robots.txt is ever consulted. Cloudflare's AI bot blocking is the most common cause. Check your CDN's bot settings. Our request does not come from the crawler's own IP addresses, so a site that verifies crawlers by IP can treat the real crawler differently.
Which AI crawlers should I allow?
If you want AI answers to name and cite you, allow the search and user-triggered crawlers: OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot and Bingbot. Training crawlers such as GPTBot, ClaudeBot, CCBot and Google-Extended are a separate choice about whether your content trains future models.