AI crawler directory
Which websites block AI crawlers?
We read the robots.txt of 1,278 popular websites to see which AI crawlers they allow, restrict or block — and whether they publish llms.txt. Updated Oct 2, 2026.
- Open to AI crawlers
- 73%
- 939 sites
- Blocks AI training, allows AI answers
- 10%
- 129 sites
- Blocks AI answer engines
- 0%
- 2 sites
- Blocks AI training and AI answers
- 16%
- 208 sites
Share of sites blocking each AI crawler
- CCBot21%
- Bytespider20%
- GPTBot19%
- ClaudeBot18%
- Google-Extended16%
- Applebot-Extended16%
- Meta-ExternalAgent16%
- PerplexityBot14%
- ChatGPT-User12%
- Claude-SearchBot10%
- Claude-User10%
- OAI-SearchBot9%
Orange: crawlers that power AI answers. Grey: crawlers for AI training. Only full blocks (Disallow: /) are counted.
19% of these sites publish an llms.txt file (245 sites).
Most popular sites
- google.comOpen to AI crawlers
- cloudflare.comOpen to AI crawlers
- facebook.comBlocks AI training and AI answers
- microsoft.comOpen to AI crawlers
- youtube.comOpen to AI crawlers
- apple.comOpen to AI crawlers
- instagram.comBlocks AI training and AI answers
- mail.ruBlocks AI training and AI answers
- twitter.comBlocks AI training and AI answers
- dzen.ruOpen to AI crawlers
- amazon.comBlocks AI training and AI answers
- github.comBlocks AI training, allows AI answers
- wikipedia.orgOpen to AI crawlers
- bing.comOpen to AI crawlers
- whatsapp.netBlocks AI training and AI answers
- netflix.comBlocks AI training, allows AI answers
- gandi.netOpen to AI crawlers
- youtu.beOpen to AI crawlers
- x.comBlocks AI training and AI answers
- wordpress.orgOpen to AI crawlers
- icloud.comOpen to AI crawlers
- pinterest.comBlocks AI training and AI answers
- tiktok.comBlocks AI training and AI answers
- yahoo.comBlocks AI training and AI answers
- roblox.comOpen to AI crawlers
- adobe.comOpen to AI crawlers
- spotify.comOpen to AI crawlers
- wa.meOpen to AI crawlers
- myfritz.netBlocks AI training and AI answers
- msn.comBlocks AI training and AI answers
- zoom.usOpen to AI crawlers
- qq.comOpen to AI crawlers
- nginx.orgOpen to AI crawlers
- baidu.comBlocks AI training and AI answers
- chatgpt.comBlocks AI training and AI answers
- opera.comOpen to AI crawlers
- vimeo.comBlocks AI training, allows AI answers
- openai.comOpen to AI crawlers
- mozilla.orgOpen to AI crawlers
- blogspot.comOpen to AI crawlers
Browse all sites A–Z
Methodology
- Sites: the most popular domains in the Tranco list (list ID Y83YG), one domain per brand, excluding infrastructure/CDN hosts and adult sites.
- We fetched each site's
/robots.txtand/llms.txtonce, identifying as AskableHQ. Sites without a readable robots.txt aren't included. - Blocked means the rules that apply to that crawler (its own group, or
*if it has none) disallow/. Some paths restricted means only parts of the site are disallowed. - robots.txt is a public request, not enforcement: sites may also block bots at the firewall, which this data doesn't capture. Crawler purposes follow each vendor's own documentation.
- Each page shows the date its data was checked. We re-check the directory regularly.