Robots.txt AI Crawler Checker
Many sites accidentally block AI crawlers with a copied robots.txt or a CDN default. Paste yours below to see exactly which bots can read your site.
Tip: open yoursite.com/robots.txt, copy everything, paste here. Parsing happens in your browser.
- Blocked
GPTBot
OpenAI model training
- Allowed
OAI-SearchBot
ChatGPT search results
- Allowed
ChatGPT-User
ChatGPT fetching a page a user asked about
- Allowed
ClaudeBot
Anthropic model training
- Allowed
Claude-SearchBot
Claude search results
- Allowed
Claude-User
Claude fetching a page a user asked about
- Allowed
anthropic-ai
Legacy Anthropic token (no longer documented)
- Allowed
PerplexityBot
Perplexity answers & citations
- Allowed
Google-Extended
Gemini training & grounding (not Google Search)
- Allowed
Applebot-Extended
Apple Intelligence
- Blocked
CCBot
Common Crawl (feeds many LLMs)
- Allowed
Bytespider
ByteDance / Doubao
- Allowed
Meta-ExternalAgent
Meta AI
User-agent: * Disallow: /admin Sitemap: https://example.com/sitemap.xml # Allow AI assistants to read and cite this site User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ClaudeBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Google-Extended Allow: /
Curious what big sites do? See which popular websites block AI crawlers, e.g. who blocks GPTBot.
Key terms: robots.txt, AI crawler
FAQ
What is GPTBot?
GPTBot is OpenAI's web crawler. OAI-SearchBot powers ChatGPT search results. Blocking them reduces the chance ChatGPT cites your site.
What is Google-Extended?
Google-Extended is a robots.txt token controlling whether your content is used for Gemini. It does not affect Google Search ranking.
Should I allow every AI crawler?
If you sell a product and want to be recommended in AI answers, allowing the major assistants is usually beneficial. Publishers protecting content may choose differently.