Does cnn.com block AI crawlers?
Blocks AI training and AI answers
cnn.com blocks both AI training crawlers and AI answer crawlers in its robots.txt.
Checked from cnn.com/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 11 of 12
- AI bots named in robots.txt
- 12
- llms.txt
- Not found
- Popularity
- Tranco top 500 (#205)
Can AI assistants cite cnn.com?
- ChatGPT search: its search crawler is blocked
- Claude: its search crawler is blocked
- Perplexity: its search crawler is blocked
cnn.com opts out of AI training by OpenAI, Anthropic, Google, Apple, Common Crawl, ByteDance. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on cnn.com |
|---|---|---|
| GPTBotOpenAI | AI training | Blocked |
| OAI-SearchBotOpenAI | AI answers | Blocked |
| ChatGPT-UserOpenAI | AI answers | Blocked |
| ClaudeBotAnthropic | AI training | Blocked |
| Claude-SearchBotAnthropic | AI answers | Blocked |
| Claude-UserAnthropic | AI answers | Blocked |
| PerplexityBotPerplexity | AI answers | Blocked |
| Google-ExtendedGoogle | AI training | Blocked |
| Applebot-ExtendedApple | AI training | Blocked |
| CCBotCommon Crawl | AI training | Blocked |
| BytespiderByteDance | AI training | Blocked |
| Meta-ExternalAgentMeta | AI training | Some paths restricted |
How cnn.com compares
- 19% of the 1,278 popular sites we checked block GPTBot; cnn.com is one of them.
- 18% of the 1,278 popular sites we checked block ClaudeBot; cnn.com is one of them.
- 14% of the 1,278 popular sites we checked block PerplexityBot; cnn.com is one of them.
- Among 177 sites of similar popularity, 34% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: AI2Bot User-agent: Ai2Bot-Dolma User-agent: AliyunSecBot User-agent: Amazonbot User-agent: amzn-searchbot User-agent: amzn-user User-agent: anthropic-ai User-agent: Applebot-Extended User-agent: Archive.org_bot User-agent: AwarioRssBot User-agent: AwarioSmartBot User-agent: Brightbot 1.0 User-agent: Bytespider User-agent: CCBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User User-agent: Claude-Web User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: Crawlspace User-agent: DataForSeoBot User-agent: Diffbot User-agent: DuckAssistBot User-agent: EchoboxBot User-agent: Exabot User-agent: FacebookBot User-agent: FriendlyCrawler User-agent: GPTBot User-agent: Google-Extended User-agent: GoogleOther User-agent: GoogleOther-Image User-agent: GoogleOther-Video User-agent: IAB-Tech-Lab User-agent: iaskspider/2.0 User-agent: ICC-Crawler User-agent: img2dataset User-agent: ISSCyberRiskCrawler User-agent: ImagesiftBot User-agent: Kangaroo Bot User-agent: magpie-crawler User-agent: MistralAI-user User-agent: MyCentralAIScraperBot User-agent: NewsNow User-agent: news-please User-agent: OAI-SearchBot User-agent: omgili User-agent: omgilibot User-agent: PanguBot User-agent: Panscient User-agent: PiplBot User-agent: PerplexityBot User-agent: Perplexity-User User-agent: PetalBot User-agent: Poseidon Research Crawler User-agent: QuillBot User-agent: quillbot.com User-agent: Quora-Bot User-agent: SBIntuitionsBot User-agent: Scrapy User-agent: SeekrBot User-agent: SemrushBot-OCOB User-agent: SemrushBot-SWA User-agent: SeznamHomepageCrawler User-agent: ShapBot User-agent: Shap-User User-agent: Sidetrade indexer bot User-agent: TaraGroup Intelligence Bot User-agent: Timpibot User-agent: TurnitinBot User-agent: VelenPublicWebCrawler User-agent: ViennaTinyBot User-agent: Webzio-Extended User-agent: YandexAdditional User-agent: YandexAdditionalBot User-agent: YouBot Disallow: / User-agent: * Allow: /partners/ipad/live-video.json Disallow: /*.jsx$ Disallow: *.jsx$ Disallow: /*.jsx/ Disallow: *.jsx? Disallow: /ads/ Disallow: /aol/ Disallow: /api/ Disallow: /beta/ Disallow: /browsers/ Disallow: /cl/ Disallow: /cnews/ Disallow: /cnn_adspaces Disallow: /cnnbeta/ Disallow: /cnnintl_adspaces Disallow: /development Disallow: /editionssi Disallow: /help/cnnx.html Disallow: /NewsPass Disallow: /NOKIA Disallow: /partners/ Disallow: /pipeline/ Disallow: /pointroll/ Disallow: /POLLSERVER/ Disallow: /pr/ Disallow: /PV/ Disallow: /Quickcast/ Disallow: /quickcast/ Disallow: /QUICKNEWS/ Disallow: /search Disallow: /subscriptions/video-docs Disallow: /terms Disallow: /test/ Disallow: /virtual/ Disallow: /WEB-INF/ Disallow: /web.projects/ Disallow: /webview/
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does cnn.com block GPTBot?
Yes. As of Oct 2, 2026, cnn.com's robots.txt disallows GPTBot, OpenAI's training crawler.
Can ChatGPT search show results from cnn.com?
Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.
Does cnn.com have an llms.txt file?
We didn't find a valid llms.txt at cnn.com/llms.txt when we checked.
Data: cnn.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.