Does newscientist.com block AI crawlers?
Blocks AI training and AI answers
newscientist.com blocks both AI training crawlers and AI answer crawlers in its robots.txt.
Checked from newscientist.com/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 12 of 12
- AI bots named in robots.txt
- 13
- llms.txt
- Not found
- Popularity
- Tranco top 5,000 (#2,773)
Can AI assistants cite newscientist.com?
- ChatGPT search: its search crawler is blocked
- Claude: its search crawler is blocked
- Perplexity: its search crawler is blocked
newscientist.com opts out of AI training by OpenAI, Anthropic, Google, Apple, Common Crawl, ByteDance, Meta. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on newscientist.com |
|---|---|---|
| GPTBotOpenAI | AI training | Blocked |
| OAI-SearchBotOpenAI | AI answers | Blocked |
| ChatGPT-UserOpenAI | AI answers | Blocked |
| ClaudeBotAnthropic | AI training | Blocked |
| Claude-SearchBotAnthropic | AI answers | Blocked |
| Claude-UserAnthropic | AI answers | Blocked |
| PerplexityBotPerplexity | AI answers | Blocked |
| Google-ExtendedGoogle | AI training | Blocked |
| Applebot-ExtendedApple | AI training | Blocked |
| CCBotCommon Crawl | AI training | Blocked |
| BytespiderByteDance | AI training | Blocked |
| Meta-ExternalAgentMeta | AI training | Blocked |
How newscientist.com compares
- 19% of the 1,278 popular sites we checked block GPTBot; newscientist.com is one of them.
- 18% of the 1,278 popular sites we checked block ClaudeBot; newscientist.com is one of them.
- 14% of the 1,278 popular sites we checked block PerplexityBot; newscientist.com is one of them.
- Among 173 sites of similar popularity, 25% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: AhrefsBot User-agent: AI2Bot User-agent: Ai2Bot-Dolma User-agent: Amazonbot User-agent: amazon-kendra- User-agent: anthropic-ai User-agent: Applebot-Extended User-agent: bedrockbot User-agent: Bytespider User-agent: CloudflareBrowserRenderingCrawler User-agent: CCBot User-agent: ChatGLM-Spider User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User User-agent: Claude-Web User-agent: cohere-ai User-agent: Cotoyogi User-agent: DeepSeekBot User-agent: Diffbot User-agent: DuckAssistBot User-agent: EchoboxBot User-agent: FacebookBot User-agent: FriendlyCrawler User-agent: Gemini-Deep-Research User-agent: Google-CloudVertexBot User-agent: Google-Extended User-agent: GoogleOther User-agent: GoogleOther-Image User-agent: GoogleOther-Video User-agent: GPTBot User-agent: Grok User-agent: iaskspider/2.0 User-agent: ICC-Crawler User-agent: ImagesiftBot User-agent: img2dataset User-agent: ISSCyberRiskCrawler User-agent: Kangaroo Bot User-agent: KunatoCrawler User-agent: LinerBot User-agent: Meltwater User-agent: Meta-ExternalAgent User-agent: Meta-ExternalFetcher User-agent: MistralAI-User User-agent: OAI-Operator User-agent: OAI-SearchBot User-agent: omgili User-agent: omgilibot User-agent: PanguBot User-agent: PerplexityBot User-agent: Perplexity-User User-agent: PetalBot User-agent: QualifiedBot User-agent: Scrapy User-agent: Seekr User-agent: Sidetrade indexer bot User-agent: TaraGroup Intelligent Bot User-agent: TikTokSpider User-agent: Timpibot User-agent: VelenPublicWebCrawler User-agent: WARDBot User-agent: Webzio-Extended User-agent: wpbot User-agent: WRTNBot User-agent: YouBot Disallow: / User-agent: * Disallow: /21632812681/ Disallow: /activate-subscription/ Disallow: /feed/ Disallow: /login/ Disallow: /logout/ Disallow: /lost-password/ Disallow: /my-account/ Disallow: /registration/ Disallow: /search/ Disallow: /upfront/ Disallow: /wp-admin/ Disallow: /api/ Disallow: /build/ Disallow: /nsj/jobsjson/ Disallow: /nsj/logon/ Disallow: /nsj/newalert/ Disallow: /nsj/analytics/ Disallow: /nsj/apply-profile/ Disallow: /nsj/document/ Disallow: /nsj/emailjob/ Disallow: /nsj/external-redirect-registration/ Disallow: /nsj/invalid-request/ Disallow: /nsj/jbequicksignup/ Disallow: /nsj/jobsrss/ Disallow: /nsj/previewjob/ Disallow: /nsj/searchjobs/ Disallow: /nsj/session-img/ Disallow: /nsj/your-jobs/ Disallow: /nsj/profile/ Disallow: /nsj/remindme/ Disallow: /nsj/searchjobs/ User-agent: * Disallow:
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does newscientist.com block GPTBot?
Yes. As of Oct 2, 2026, newscientist.com's robots.txt disallows GPTBot, OpenAI's training crawler.
Can ChatGPT search show results from newscientist.com?
Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.
Does newscientist.com have an llms.txt file?
We didn't find a valid llms.txt at newscientist.com/llms.txt when we checked.
Data: newscientist.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.