Does irishtimes.com block AI crawlers?
Blocks AI training and AI answers
irishtimes.com blocks both AI training crawlers and AI answer crawlers in its robots.txt.
Checked from irishtimes.com/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 2 of 12
- AI bots named in robots.txt
- 2
- llms.txt
- Not found
- Popularity
- Tranco top 5,000 (#2,997)
Can AI assistants cite irishtimes.com?
- ChatGPT search: its search crawler is blocked
- Claude: its search crawlers may read the site
- Perplexity: its search crawlers may read the site
irishtimes.com opts out of AI training by OpenAI. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on irishtimes.com |
|---|---|---|
| GPTBotOpenAI | AI training | Blocked |
| OAI-SearchBotOpenAI | AI answers | Some paths restricted |
| ChatGPT-UserOpenAI | AI answers | Blocked |
| ClaudeBotAnthropic | AI training | Some paths restricted |
| Claude-SearchBotAnthropic | AI answers | Some paths restricted |
| Claude-UserAnthropic | AI answers | Some paths restricted |
| PerplexityBotPerplexity | AI answers | Some paths restricted |
| Google-ExtendedGoogle | AI training | Some paths restricted |
| Applebot-ExtendedApple | AI training | Some paths restricted |
| CCBotCommon Crawl | AI training | Some paths restricted |
| BytespiderByteDance | AI training | Some paths restricted |
| Meta-ExternalAgentMeta | AI training | Some paths restricted |
How irishtimes.com compares
- 19% of the 1,278 popular sites we checked block GPTBot; irishtimes.com is one of them.
- 18% of the 1,278 popular sites we checked block ClaudeBot; irishtimes.com does not.
- 14% of the 1,278 popular sites we checked block PerplexityBot; irishtimes.com does not.
- Among 169 sites of similar popularity, 25% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: ChatGPT-User Disallow: / User-agent: GPTBot Disallow: / User-agent: * Disallow: /blogimageupload/ Disallow: /captcha/ Disallow: /content/ Disallow: /cmlink/ Disallow: /cm/ Disallow: /error/ Disallow: /errorpages/ Disallow: /logger/ Disallow: /mailapi/ Disallow: /membership/ Disallow: /mobile/ Disallow: /newspaper/archive/ Disallow: /photosales/index.cfm?fuseaction= Disallow: /poll/ Disallow: /polopolydevelopment/ Disallow: /redirect/ Disallow: /rta-logging/reader-history.php Disallow: /search/ Disallow: /search-results/ Disallow: /search/archive.html Disallow: /search/search-7.4195619 Disallow: /search/search-7.1213540 Disallow: /search/search-7.2285082 Disallow: /status/ Disallow: /zephr/feature-decisions Disallow: /zephr/features Disallow: /polopoly_fs/ Allow: /subscribe/core/*.css$ Allow: /subscribe/core/*.css? Allow: /subscribe/core/*.js$ Allow: /subscribe/core/*.js? Allow: /subscribe/core/*.gif Allow: /subscribe/core/*.jpg Allow: /subscribe/core/*.jpeg Allow: /subscribe/core/*.png Allow: /subscribe/core/*.svg Allow: /subscribe/profiles/*.css$ Allow: /subscribe/profiles/*.css? Allow: /subscribe/profiles/*.js$ Allow: /subscribe/profiles/*.js?
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does irishtimes.com block GPTBot?
Yes. As of Oct 2, 2026, irishtimes.com's robots.txt disallows GPTBot, OpenAI's training crawler.
Can ChatGPT search show results from irishtimes.com?
Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.
Does irishtimes.com have an llms.txt file?
We didn't find a valid llms.txt at irishtimes.com/llms.txt when we checked.
Data: irishtimes.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.