Does mirror.co.uk block AI crawlers?
Blocks AI training and AI answers
mirror.co.uk blocks both AI training crawlers and AI answer crawlers in its robots.txt.
Checked from mirror.co.uk/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 7 of 12
- AI bots named in robots.txt
- 8
- llms.txt
- Not found
- Popularity
- Tranco top 5,000 (#2,261)
Can AI assistants cite mirror.co.uk?
- ChatGPT search: its search crawler is blocked
- Claude: its search crawlers may read the site
- Perplexity: its search crawler is blocked
mirror.co.uk opts out of AI training by OpenAI, Anthropic, Apple, Common Crawl, Meta. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on mirror.co.uk |
|---|---|---|
| GPTBotOpenAI | AI training | Blocked |
| OAI-SearchBotOpenAI | AI answers | Blocked |
| ChatGPT-UserOpenAI | AI answers | Some paths restricted |
| ClaudeBotAnthropic | AI training | Blocked |
| Claude-SearchBotAnthropic | AI answers | Some paths restricted |
| Claude-UserAnthropic | AI answers | Some paths restricted |
| PerplexityBotPerplexity | AI answers | Blocked |
| Google-ExtendedGoogle | AI training | Some paths restricted |
| Applebot-ExtendedApple | AI training | Blocked |
| CCBotCommon Crawl | AI training | Blocked |
| BytespiderByteDance | AI training | Some paths restricted |
| Meta-ExternalAgentMeta | AI training | Blocked |
How mirror.co.uk compares
- 19% of the 1,278 popular sites we checked block GPTBot; mirror.co.uk is one of them.
- 18% of the 1,278 popular sites we checked block ClaudeBot; mirror.co.uk is one of them.
- 14% of the 1,278 popular sites we checked block PerplexityBot; mirror.co.uk is one of them.
- Among 187 sites of similar popularity, 21% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: * Disallow: /topics/* Disallow: /topics Disallow: /*token=* Disallow: /search/ Disallow: /comm-part-test/ Disallow: /*service=ajax Disallow: /centenary-fund/ Disallow: /3am/weird-celeb-news/xxx-1341448 Disallow: /3am/weird-celeb-news/tamara-ecclestone-watched-boyfriends-sex-1341448 Disallow: /comm-part-test/oh-like-beside-seaside-uk-1722203 Disallow: /resources/js/s_code.js Disallow: /template/ Disallow: /tv/tv-news/jean-alexander-dies-coronation-streets-3764827 Disallow: /3am/celebrity-news/anne-kirkbride-dead-bill-roache-5013775 Disallow: /lifestyle/cartoons/andy-capp/andy-capp Disallow: /lifestyle/cartoons/the-gag-vault/gag-vault Disallow: /lifestyle/cartoons/perishers/perishers Disallow: /lifestyle/cartoons/horace/horace Disallow: /lifestyle/cartoons/garth/garth Disallow: /lifestyle/cartoons/mandy/mandy Disallow: /lifestyle/cartoons/kerber-black/ Disallow: /regression-test-home/ Disallow: /5293/ Disallow: /exclusive-offers/ User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: PerplexityBot Disallow: / User-agent: CCBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Meta-ExternalAgent Disallow: /
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does mirror.co.uk block GPTBot?
Yes. As of Oct 2, 2026, mirror.co.uk's robots.txt disallows GPTBot, OpenAI's training crawler.
Can ChatGPT search show results from mirror.co.uk?
Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.
Does mirror.co.uk have an llms.txt file?
We didn't find a valid llms.txt at mirror.co.uk/llms.txt when we checked.
Data: mirror.co.uk's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.