Skip to content
AskableHQ
Launch

Does oreilly.com block AI crawlers?

Blocks AI training, allows AI answers

oreilly.com blocks at least one AI training crawler but still lets AI answer engines read its pages.

Checked from oreilly.com/robots.txt. Robots rules change; the live file is the source of truth.

AI crawlers blocked
4 of 12
AI bots named in robots.txt
12
llms.txt
Published
Popularity
Tranco top 2,000 (#1,941)

Can AI assistants cite oreilly.com?

  • ChatGPT search: its search crawlers may read the site
  • Claude: its search crawlers may read the site
  • Perplexity: its search crawlers may read the site

oreilly.com opts out of AI training by OpenAI, Common Crawl, ByteDance, Meta. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.

AI crawler access, bot by bot

CrawlerUsed forStatus on oreilly.com
GPTBotOpenAIAI trainingBlocked
OAI-SearchBotOpenAIAI answersSome paths restricted
ChatGPT-UserOpenAIAI answersSome paths restricted
ClaudeBotAnthropicAI trainingSome paths restricted
Claude-SearchBotAnthropicAI answersSome paths restricted
Claude-UserAnthropicAI answersSome paths restricted
PerplexityBotPerplexityAI answersSome paths restricted
Google-ExtendedGoogleAI trainingSome paths restricted
Applebot-ExtendedAppleAI trainingSome paths restricted
CCBotCommon CrawlAI trainingBlocked
BytespiderByteDanceAI trainingBlocked
Meta-ExternalAgentMetaAI trainingBlocked

How oreilly.com compares

  • 19% of the 1,278 popular sites we checked block GPTBot; oreilly.com is one of them.
  • 18% of the 1,278 popular sites we checked block ClaudeBot; oreilly.com does not.
  • 14% of the 1,278 popular sites we checked block PerplexityBot; oreilly.com does not.
  • Among 158 sites of similar popularity, 25% block at least one AI crawler.

The robots.txt rules that apply to AI crawlers

User-agent: *
Disallow: /images/
Disallow: /graphics/
Disallow: /admin/
Disallow: /promos/
Disallow: /ddp/
Disallow: /dpp/
Disallow: /programming/free/files/
Disallow: /design/free/files/
Disallow: /iot/free/files/
Disallow: /data/free/files/
Disallow: /webops-perf/free/files/
Disallow: /web-platform/free/files/
Disallow: /cs/
Disallow: /test/
Disallow: /*/?ar
Disallow: /*/?orpq
Disallow: /*/?discount=learn
Disallow: /self-registration/*
Disallow: /member/login/*?next=

User-agent: ChatGPT-User
User-agent: ChatGPT-User/2.0
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: DuckAssistBot
User-agent: Applebot
User-agent: bingbot
User-agent: Googlebot
Allow: /
Disallow: /member/login/*?next=

User-agent: ClaudeBot
User-agent: claude-web
User-agent: Claude-SearchBot
User-agent: YouBot
User-agent: Applebot-Extended
User-agent: MistralAI-User
User-agent: meta-webindexer
User-agent: GoogleAgent-Mariner
User-agent: Google-Extended
User-agent: DotBot
User-agent: archive.org_bot
User-agent: Pinterestbot
User-agent: TelegramBot
User-agent: kakaotalk-scrap
User-agent: getstream.io/opengraph-bot
User-agent: RamblerMail
User-agent: Brave
User-agent: OAI-SearchBot
Allow: /
Disallow: /member/login/*?next=

User-agent: GPTBot
User-agent: anthropic-ai
User-agent: cohere-ai
User-agent: CCBot
User-agent: AI2Bot
User-agent: Amazonbot
User-agent: Amazonbot-Video
User-agent: Bytespider
User-agent: meta-externalagent
User-agent: Diffbot
User-agent: omgili
User-agent: TimpiBot
User-agent: SeznamBot
User-agent: Exabot
User-agent: YandexBot
User-agent: Sogou
User-agent: 360Spider
User-agent: YisouSpider
User-agent: Baiduspider-Render/2.0
User-agent: MJ12bot
User-agent: DataForSeoBot
User-agent: AliyunSecBot
User-agent: ArchiveTeam
User-agent: ArchiveTeam ArchiveBot
User-agent: AwarioBot
User-agent: ZoominfoBot
User-agent: ali-implementer
User-agent: Blueno
User-agent: BIGO-baiguoyuan
User-agent: WorksOgCrawler
User-agent: OnPageBot
User-agent: DoCoMo
User-agent: SAMSUNG-SGH-E250
User-agent: TA-Googlebot
Disallow: /

Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.

Sites with a similar AI policy

All sites that are “blocks ai training, allows ai answers” →

Does oreilly.com block GPTBot?

Yes. As of Oct 2, 2026, oreilly.com's robots.txt disallows GPTBot, OpenAI's training crawler.

Can ChatGPT search show results from oreilly.com?

Its robots.txt doesn't block OAI-SearchBot or ChatGPT-User, so ChatGPT's search crawlers may read it. Whether pages are actually shown depends on ChatGPT.

Does oreilly.com have an llms.txt file?

Yes — oreilly.com/llms.txt was found when we checked.

Data: oreilly.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.