Skip to content
AskableHQ
Launch

Does cnn.com block AI crawlers?

Blocks AI training and AI answers

cnn.com blocks both AI training crawlers and AI answer crawlers in its robots.txt.

Checked from cnn.com/robots.txt. Robots rules change; the live file is the source of truth.

AI crawlers blocked
11 of 12
AI bots named in robots.txt
12
llms.txt
Not found
Popularity
Tranco top 500 (#205)

Can AI assistants cite cnn.com?

  • ChatGPT search: its search crawler is blocked
  • Claude: its search crawler is blocked
  • Perplexity: its search crawler is blocked

cnn.com opts out of AI training by OpenAI, Anthropic, Google, Apple, Common Crawl, ByteDance. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.

AI crawler access, bot by bot

CrawlerUsed forStatus on cnn.com
GPTBotOpenAIAI trainingBlocked
OAI-SearchBotOpenAIAI answersBlocked
ChatGPT-UserOpenAIAI answersBlocked
ClaudeBotAnthropicAI trainingBlocked
Claude-SearchBotAnthropicAI answersBlocked
Claude-UserAnthropicAI answersBlocked
PerplexityBotPerplexityAI answersBlocked
Google-ExtendedGoogleAI trainingBlocked
Applebot-ExtendedAppleAI trainingBlocked
CCBotCommon CrawlAI trainingBlocked
BytespiderByteDanceAI trainingBlocked
Meta-ExternalAgentMetaAI trainingSome paths restricted

How cnn.com compares

  • 19% of the 1,278 popular sites we checked block GPTBot; cnn.com is one of them.
  • 18% of the 1,278 popular sites we checked block ClaudeBot; cnn.com is one of them.
  • 14% of the 1,278 popular sites we checked block PerplexityBot; cnn.com is one of them.
  • Among 177 sites of similar popularity, 34% block at least one AI crawler.

The robots.txt rules that apply to AI crawlers

User-agent: AI2Bot
User-agent: Ai2Bot-Dolma
User-agent: AliyunSecBot
User-agent: Amazonbot
User-agent: amzn-searchbot
User-agent: amzn-user
User-agent: anthropic-ai
User-agent: Applebot-Extended
User-agent: Archive.org_bot
User-agent: AwarioRssBot
User-agent: AwarioSmartBot
User-agent: Brightbot 1.0
User-agent: Bytespider
User-agent: CCBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: Claude-Web
User-agent: cohere-ai
User-agent: cohere-training-data-crawler
User-agent: Crawlspace
User-agent: DataForSeoBot
User-agent: Diffbot
User-agent: DuckAssistBot
User-agent: EchoboxBot
User-agent: Exabot
User-agent: FacebookBot
User-agent: FriendlyCrawler
User-agent: GPTBot
User-agent: Google-Extended
User-agent: GoogleOther
User-agent: GoogleOther-Image
User-agent: GoogleOther-Video
User-agent: IAB-Tech-Lab
User-agent: iaskspider/2.0
User-agent: ICC-Crawler
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: ImagesiftBot
User-agent: Kangaroo Bot
User-agent: magpie-crawler
User-agent: MistralAI-user
User-agent: MyCentralAIScraperBot
User-agent: NewsNow
User-agent: news-please
User-agent: OAI-SearchBot
User-agent: omgili
User-agent: omgilibot
User-agent: PanguBot
User-agent: Panscient
User-agent: PiplBot
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: PetalBot
User-agent: Poseidon Research Crawler
User-agent: QuillBot
User-agent: quillbot.com
User-agent: Quora-Bot
User-agent: SBIntuitionsBot
User-agent: Scrapy
User-agent: SeekrBot
User-agent: SemrushBot-OCOB
User-agent: SemrushBot-SWA
User-agent: SeznamHomepageCrawler
User-agent: ShapBot
User-agent: Shap-User
User-agent: Sidetrade indexer bot
User-agent: TaraGroup Intelligence Bot
User-agent: Timpibot
User-agent: TurnitinBot
User-agent: VelenPublicWebCrawler
User-agent: ViennaTinyBot
User-agent: Webzio-Extended
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
User-agent: YouBot
Disallow: /

User-agent: *
Allow: /partners/ipad/live-video.json
Disallow: /*.jsx$
Disallow: *.jsx$
Disallow: /*.jsx/
Disallow: *.jsx?
Disallow: /ads/
Disallow: /aol/
Disallow: /api/
Disallow: /beta/
Disallow: /browsers/
Disallow: /cl/
Disallow: /cnews/
Disallow: /cnn_adspaces
Disallow: /cnnbeta/
Disallow: /cnnintl_adspaces
Disallow: /development
Disallow: /editionssi
Disallow: /help/cnnx.html
Disallow: /NewsPass
Disallow: /NOKIA
Disallow: /partners/
Disallow: /pipeline/
Disallow: /pointroll/
Disallow: /POLLSERVER/
Disallow: /pr/
Disallow: /PV/
Disallow: /Quickcast/
Disallow: /quickcast/
Disallow: /QUICKNEWS/
Disallow: /search
Disallow: /subscriptions/video-docs
Disallow: /terms
Disallow: /test/
Disallow: /virtual/
Disallow: /WEB-INF/
Disallow: /web.projects/
Disallow: /webview/

Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.

Sites with a similar AI policy

All sites that are “blocks ai training and ai answers” →

Does cnn.com block GPTBot?

Yes. As of Oct 2, 2026, cnn.com's robots.txt disallows GPTBot, OpenAI's training crawler.

Can ChatGPT search show results from cnn.com?

Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.

Does cnn.com have an llms.txt file?

We didn't find a valid llms.txt at cnn.com/llms.txt when we checked.

Data: cnn.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.