Does cambridge.org block AI crawlers?
Blocks AI answer engines
cambridge.org blocks at least one AI answer crawler, which can keep it out of those assistants' cited answers.
Checked from cambridge.org/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 1 of 12
- AI bots named in robots.txt
- 1
- llms.txt
- Not found
- Popularity
- Tranco top 1,000 (#680)
Can AI assistants cite cambridge.org?
- ChatGPT search: its search crawler is blocked
- Claude: its search crawlers may read the site
- Perplexity: its search crawlers may read the site
cambridge.org doesn't opt out of AI model training through robots.txt. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on cambridge.org |
|---|---|---|
| GPTBotOpenAI | AI training | Some paths restricted |
| OAI-SearchBotOpenAI | AI answers | Some paths restricted |
| ChatGPT-UserOpenAI | AI answers | Blocked |
| ClaudeBotAnthropic | AI training | Some paths restricted |
| Claude-SearchBotAnthropic | AI answers | Some paths restricted |
| Claude-UserAnthropic | AI answers | Some paths restricted |
| PerplexityBotPerplexity | AI answers | Some paths restricted |
| Google-ExtendedGoogle | AI training | Some paths restricted |
| Applebot-ExtendedApple | AI training | Some paths restricted |
| CCBotCommon Crawl | AI training | Some paths restricted |
| BytespiderByteDance | AI training | Some paths restricted |
| Meta-ExternalAgentMeta | AI training | Some paths restricted |
How cambridge.org compares
- 19% of the 1,278 popular sites we checked block GPTBot; cambridge.org does not.
- 18% of the 1,278 popular sites we checked block ClaudeBot; cambridge.org does not.
- 14% of the 1,278 popular sites we checked block PerplexityBot; cambridge.org does not.
- Among 194 sites of similar popularity, 27% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: ChatGPT-User Disallow: / User-agent: * Allow: /concrete/*.css Allow: /concrete/*.js Allow: /concrete/*.jpg Allow: /concrete/*.svg Allow: /concrete/*.png Allow: /concrete/*.gif Allow: /packages/*.css Allow: /packages/*.js Allow: /packages/*.jpg Allow: /packages/*.svg Allow: /packages/*.png Allow: /packages/*.gif Allow: /blocks/*.css Allow: /blocks/*.js Allow: /blocks/*.jpg Allow: /blocks/*.svg Allow: /blocks/*.png Allow: /blocks/*.gif Allow: /tools/packages/cambridge_themes/getSocialImage Allow: /tools/packages/cambridge_themes/miniCart* Disallow: /aca/authorinformation/ Disallow: /blocks Disallow: /concrete Disallow: /config Disallow: /controllers Disallow: /css Disallow: /elements Disallow: /helpers Disallow: /jobs Disallow: /js Disallow: /languages Disallow: /libraries Disallow: /mail Disallow: /models Disallow: /packages Disallow: /single_pages Disallow: /themes Disallow: /tools Disallow: /updates Disallow: /branding
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does cambridge.org block GPTBot?
No. As of Oct 2, 2026, cambridge.org's robots.txt doesn't block GPTBot, though some paths are restricted for all crawlers.
Can ChatGPT search show results from cambridge.org?
Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.
Does cambridge.org have an llms.txt file?
We didn't find a valid llms.txt at cambridge.org/llms.txt when we checked.
Data: cambridge.org's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.