Skip to content
AskableHQ
Launch

Does reuters.com block AI crawlers?

Blocks AI training and AI answers

reuters.com blocks both AI training crawlers and AI answer crawlers in its robots.txt.

Checked from reuters.com/robots.txt. Robots rules change; the live file is the source of truth.

AI crawlers blocked
10 of 12
AI bots named in robots.txt
3
llms.txt
Not found
Popularity
Tranco top 500 (#291)

Can AI assistants cite reuters.com?

  • ChatGPT search: its search crawlers may read the site
  • Claude: its search crawler is blocked
  • Perplexity: its search crawler is blocked

reuters.com opts out of AI training by OpenAI, Anthropic, Google, Apple, Common Crawl, ByteDance, Meta. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.

AI crawler access, bot by bot

CrawlerUsed forStatus on reuters.com
GPTBotOpenAIAI trainingBlocked
OAI-SearchBotOpenAIAI answersSome paths restricted
ChatGPT-UserOpenAIAI answersSome paths restricted
ClaudeBotAnthropicAI trainingBlocked
Claude-SearchBotAnthropicAI answersBlocked
Claude-UserAnthropicAI answersBlocked
PerplexityBotPerplexityAI answersBlocked
Google-ExtendedGoogleAI trainingBlocked
Applebot-ExtendedAppleAI trainingBlocked
CCBotCommon CrawlAI trainingBlocked
BytespiderByteDanceAI trainingBlocked
Meta-ExternalAgentMetaAI trainingBlocked

How reuters.com compares

  • 19% of the 1,278 popular sites we checked block GPTBot; reuters.com is one of them.
  • 18% of the 1,278 popular sites we checked block ClaudeBot; reuters.com is one of them.
  • 14% of the 1,278 popular sites we checked block PerplexityBot; reuters.com is one of them.
  • Among 192 sites of similar popularity, 32% block at least one AI crawler.

The robots.txt rules that apply to AI crawlers

User-agent: AASA-Bot
User-agent: ADmantX
User-agent: AdmiralBot
User-agent: AdsBot-Google
User-agent: AdsBot-Google-Mobile
User-agent: AmazonAdBot
User-agent: Amzn-User
User-agent: Applebot
User-agent: AppleNewsBot
User-agent: Bingbot
User-agent: BingPreview
User-agent: BrightEdge
User-agent: BrightEdgeOnCrawl
User-agent: CensysInspect
User-agent: ChatGPT-User
User-agent: Cision
User-agent: Clickagy
User-agent: Concert
User-agent: ContextualBot
User-agent: CriteoBot
User-agent: DatadogSynthetics
User-agent: datadome-pageprotect-scanner
User-agent: Discordbot
User-agent: doubleverify
User-agent: ElevenlabsBot
User-agent: Embedly
User-agent: facebookexternalhit
User-agent: Googlebot
User-agent: Googlebot Smartphone
User-agent: Googlebot-News
User-agent: Google-Display-Ads-Bot
User-agent: Google-InspectionTool
User-agent: GoogleOther
User-agent: Google-Read-Aloud
User-agent: Google-Safety
User-agent: Google-Site-Verification
User-agent: GTmetrix
User-agent: GumGumBot
User-agent: ias_crawler
User-agent: Iframely
User-agent: leiki
User-agent: LinkedInBot
User-agent: LinkTiger
User-agent: Mantisbot
User-agent: Mediapartners-Google
User-agent: meta-externalads
User-agent: meta-webindexer
User-agent: MicrosoftPreview
User-agent: MJ12bot
User-agent: Moreover
User-agent: msnbot
User-agent: NFBNewslineRobot
User-agent: OAI-SearchBot
User-agent: Oncrawl
User-agent: Opebot-v
User-agent: Opoint
User-agent: Optimizer
User-agent: outbrain
User-agent: Pinterestbot
User-agent: Prerender
User-agent: Proofpoint
User-agent: proximic
User-agent: PubMatic Crawler Bot
User-agent: Quantcastbot
User-agent: Qwantbot
User-agent: Reuters SEO Screaming Frog Spider 007
User-agent: Reuters-NAUWI
User-agent: Scom-Crawler-For-Reuters
User-agent: SinceraSyntheticUser
User-agent: Slurp
User-agent: SmartologyBot
User-agent: snews
User-agent: SocialFlow
User-agent: StatusCake
User-agent: Storebot-Google
User-agent: Stripebot
User-agent: TTD-Content
User-agent: Twitterbot
User-agent: URLDefense
User-agent: Verity
User-agent: vuln_scan_by_trustedsite_com_halo_security
User-agent: WISEbot
User-agent: Xenu Link Sleuth
User-agent: Yahoo Link Preview
User-agent: Yahoo! JAPAN
User-agent: YahooMailProxy
Disallow: /finance/stocks/option
Disallow: /finance/stocks/financialHighlights
Disallow: /search
Disallow: /site-search/
Disallow: /beta
Disallow: /designtech
Disallow: /featured-optimize
Disallow: /energy-test
Disallow: /article/beta
Disallow: /sponsored/previewcampaign
Disallow: /sponsored/previewarticle
Disallow: /test/
Disallow: /news/archive/commentary
Disallow: /brandfeatures/venture-capital
Disallow: /assets/siteindex
Disallow: /article/api/
Disallow: /practical-law-the-journal/search/
Disallow: /pf/api/
Disallow: /fr/
Disallow: /it/
Disallow: /es/
Disallow: /pt/
Disallow: /de/
Disallow: /latam/
Disallow: /account/subscribe/payment
Disallow: /*/site-search/

User-agent: Applebot-Extended
Disallow: /

User-agent: *
Allow: /plus/
Disallow: /

Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.

Sites with a similar AI policy

All sites that are “blocks ai training and ai answers” →

Does reuters.com block GPTBot?

Yes. As of Oct 2, 2026, reuters.com's robots.txt disallows GPTBot, OpenAI's training crawler.

Can ChatGPT search show results from reuters.com?

Its robots.txt doesn't block OAI-SearchBot or ChatGPT-User, so ChatGPT's search crawlers may read it. Whether pages are actually shown depends on ChatGPT.

Does reuters.com have an llms.txt file?

We didn't find a valid llms.txt at reuters.com/llms.txt when we checked.

Data: reuters.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.