Does ubuntu.com block AI crawlers?
Open to AI crawlers
ubuntu.com doesn't block any of the major AI crawlers we track in its robots.txt.
Checked from ubuntu.com/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 0 of 12
- AI bots named in robots.txt
- 11
- llms.txt
- Published
- Popularity
- Tranco top 500 (#220)
Can AI assistants cite ubuntu.com?
- ChatGPT search: its search crawlers may read the site
- Claude: its search crawlers may read the site
- Perplexity: its search crawlers may read the site
ubuntu.com doesn't opt out of AI model training through robots.txt. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on ubuntu.com |
|---|---|---|
| GPTBotOpenAI | AI training | Allowed |
| OAI-SearchBotOpenAI | AI answers | Allowed |
| ChatGPT-UserOpenAI | AI answers | Allowed |
| ClaudeBotAnthropic | AI training | Allowed |
| Claude-SearchBotAnthropic | AI answers | Allowed |
| Claude-UserAnthropic | AI answers | Some paths restricted |
| PerplexityBotPerplexity | AI answers | Allowed |
| Google-ExtendedGoogle | AI training | Allowed |
| Applebot-ExtendedApple | AI training | Allowed |
| CCBotCommon Crawl | AI training | Some paths restricted |
| BytespiderByteDance | AI training | Allowed |
| Meta-ExternalAgentMeta | AI training | Allowed |
How ubuntu.com compares
- 19% of the 1,278 popular sites we checked block GPTBot; ubuntu.com does not.
- 18% of the 1,278 popular sites we checked block ClaudeBot; ubuntu.com does not.
- 14% of the 1,278 popular sites we checked block PerplexityBot; ubuntu.com does not.
- Among 184 sites of similar popularity, 34% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: * Disallow: /search Disallow: /search* Disallow: /*/search* Disallow: /account Disallow: /account/* Disallow: /login Disallow: /logout Disallow: /pro/dashboard Disallow: /pro/users Disallow: /pro/account-users Disallow: /pro/subscribe Disallow: /pro/activate Disallow: /pro/attach Disallow: /pro/offer Disallow: /pro/offers Disallow: /pro/renewals/ Disallow: /pro/contracts/ Disallow: /pro/trial/ Disallow: /pro/set-auto-renewal Disallow: /pro/user-subscriptions Disallow: /pro/distributor/users Disallow: /pro/distributor/invoice Disallow: /pro/distributor/thank-you Disallow: /account.json Disallow: /mirrors.json Disallow: /pro/subscriptions.json Disallow: /pro/offers.json Disallow: /pro/channel-offers.json Disallow: /thank-you Disallow: /*/thank-you Disallow: /blog/draft-blogs Disallow: /blog/draft-blogs/* Disallow: /tests/ Disallow: /tests/* Disallow: /sentry-debug Disallow: /mobile Disallow: /mobile/* Disallow: /phone Disallow: /phone/* Disallow: /tablet User-agent: GPTBot User-agent: ChatGPT-User User-agent: OAI-SearchBot User-agent: ClaudeBot User-agent: Claude-Web User-agent: anthropic-ai User-agent: Claude-SearchBot User-agent: Google-Extended User-agent: meta-externalagent User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Applebot-Extended User-agent: cohere-ai User-agent: Bytespider Allow: /*?format=md Allow: /server Allow: /desktop Allow: /cloud Allow: /openstack Allow: /kubernetes Allow: /ceph Allow: /containers Allow: /core Allow: /ai Allow: /pro Allow: /landscape Allow: /security Allow: /internet-of-things Allow: /embedded Allow: /hpc Allow: /real-time Allow: /confidential-computing Allow: /enterprise-store Allow: /kernel Allow: /toolchains Allow: /robotics Allow: /certified Allow: /about Allow: /community Allow: /download Allow: /pricing Allow: /training Allow: /credentials Allow: /support Allow: /managed Allow: /managed-infrastructure Allow: /aws Allow: /azure Allow: /gcp Allow: /dell Allow: /ibm Allow: /nvidia Allow: /hpe Allow: /supermicro
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does ubuntu.com block GPTBot?
No. As of Oct 2, 2026, ubuntu.com's robots.txt doesn't block GPTBot.
Can ChatGPT search show results from ubuntu.com?
Its robots.txt doesn't block OAI-SearchBot or ChatGPT-User, so ChatGPT's search crawlers may read it. Whether pages are actually shown depends on ChatGPT.
Does ubuntu.com have an llms.txt file?
Yes — ubuntu.com/llms.txt was found when we checked.
Data: ubuntu.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.