Does arxiv.org block AI crawlers?
Open to AI crawlers
arxiv.org doesn't block any of the major AI crawlers we track in its robots.txt.
Checked from arxiv.org/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 0 of 12
- AI bots named in robots.txt
- 0
- llms.txt
- Not found
- Popularity
- Tranco top 500 (#464)
Can AI assistants cite arxiv.org?
- ChatGPT search: its search crawlers may read the site
- Claude: its search crawlers may read the site
- Perplexity: its search crawlers may read the site
arxiv.org doesn't opt out of AI model training through robots.txt. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on arxiv.org |
|---|---|---|
| GPTBotOpenAI | AI training | Some paths restricted |
| OAI-SearchBotOpenAI | AI answers | Some paths restricted |
| ChatGPT-UserOpenAI | AI answers | Some paths restricted |
| ClaudeBotAnthropic | AI training | Some paths restricted |
| Claude-SearchBotAnthropic | AI answers | Some paths restricted |
| Claude-UserAnthropic | AI answers | Some paths restricted |
| PerplexityBotPerplexity | AI answers | Some paths restricted |
| Google-ExtendedGoogle | AI training | Some paths restricted |
| Applebot-ExtendedApple | AI training | Some paths restricted |
| CCBotCommon Crawl | AI training | Some paths restricted |
| BytespiderByteDance | AI training | Some paths restricted |
| Meta-ExternalAgentMeta | AI training | Some paths restricted |
How arxiv.org compares
- 19% of the 1,278 popular sites we checked block GPTBot; arxiv.org does not.
- 18% of the 1,278 popular sites we checked block ClaudeBot; arxiv.org does not.
- 14% of the 1,278 popular sites we checked block PerplexityBot; arxiv.org does not.
- Among 208 sites of similar popularity, 29% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: * Allow: /archive Allow: /year Allow: /list Allow: /abs Allow: /pdf Allow: /html Allow: /catchup Disallow: /user Disallow: /e-print Disallow: /src Disallow: /ps Disallow: /dvi Disallow: /cookies Disallow: /form Disallow: /find Disallow: /view Disallow: /ftp Disallow: /refs Disallow: /cits Disallow: /format Disallow: /PS_cache Disallow: /Stats Disallow: /seek-and-destroy Disallow: /IgnoreMe Disallow: /oai2 Disallow: /auth Disallow: /tb Disallow: /tb-recent Disallow: /trackback Disallow: /prevnext Disallow: /ct Disallow: /api Disallow: /search Disallow: /set_author_id Disallow: /show-email
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does arxiv.org block GPTBot?
No. As of Oct 2, 2026, arxiv.org's robots.txt doesn't block GPTBot, though some paths are restricted for all crawlers.
Can ChatGPT search show results from arxiv.org?
Its robots.txt doesn't block OAI-SearchBot or ChatGPT-User, so ChatGPT's search crawlers may read it. Whether pages are actually shown depends on ChatGPT.
Does arxiv.org have an llms.txt file?
We didn't find a valid llms.txt at arxiv.org/llms.txt when we checked.
Data: arxiv.org's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.