Does nytimes.com block AI crawlers?
Blocks AI training and AI answers
nytimes.com blocks both AI training crawlers and AI answer crawlers in its robots.txt.
Checked from nytimes.com/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 12 of 12
- AI bots named in robots.txt
- 13
- llms.txt
- Not found
- Popularity
- Tranco top 500 (#157)
Can AI assistants cite nytimes.com?
- ChatGPT search: its search crawler is blocked
- Claude: its search crawler is blocked
- Perplexity: its search crawler is blocked
nytimes.com opts out of AI training by OpenAI, Anthropic, Google, Apple, Common Crawl, ByteDance, Meta. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on nytimes.com |
|---|---|---|
| GPTBotOpenAI | AI training | Blocked |
| OAI-SearchBotOpenAI | AI answers | Blocked |
| ChatGPT-UserOpenAI | AI answers | Blocked |
| ClaudeBotAnthropic | AI training | Blocked |
| Claude-SearchBotAnthropic | AI answers | Blocked |
| Claude-UserAnthropic | AI answers | Blocked |
| PerplexityBotPerplexity | AI answers | Blocked |
| Google-ExtendedGoogle | AI training | Blocked |
| Applebot-ExtendedApple | AI training | Blocked |
| CCBotCommon Crawl | AI training | Blocked |
| BytespiderByteDance | AI training | Blocked |
| Meta-ExternalAgentMeta | AI training | Blocked |
How nytimes.com compares
- 19% of the 1,278 popular sites we checked block GPTBot; nytimes.com is one of them.
- 18% of the 1,278 popular sites we checked block ClaudeBot; nytimes.com is one of them.
- 14% of the 1,278 popular sites we checked block PerplexityBot; nytimes.com is one of them.
- Among 155 sites of similar popularity, 35% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: * User-agent: Googlebot Disallow: /ads/ Disallow: /adx/bin/ Disallow: /athletic/wp/wp-admin/ Allow: /athletic/wp/wp-admin/admin-ajax.php Disallow: /athletic/async-* Disallow: /athletic/search/* Allow: /athletic/search/$ Disallow: /athletic/checkout/ Disallow: /athletic/checkout?plan_id* Allow: /athletic/checkout/$ Disallow: /athletic/checkout2* Disallow: /athletic/login/ Disallow: /athletic/login?login_source* Disallow: /athletic/login?ref_page* Allow: /athletic/login/$ Disallow: /athletic/login2/ Disallow: /athletic/login2?login_source* Disallow: /athletic/login2?ref_page* Allow: /athletic/login2/$ Disallow: /athletic/report/ Disallow: /athletic/*/discuss/* Allow: /athletic/live-blogs/discuss/ Allow: /athletic/mlb/game/discuss/ Disallow: /athletic/register/ Disallow: /athletic/register?welcome_redirect* Disallow: /athletic/register2/ Disallow: /athletic/register2?welcome_redirect* Disallow: /athletic/betmgm-redirect* Disallow: /athletic/cdn-cgi/ Disallow: /athletic/verizon/* Disallow: /athletic/forgot-password/* Disallow: /athletic/forgot-password2/* Disallow: /athletic/amp-social-login* Disallow: /athletic/track-analytics/ Disallow: /athletic/amp-auth/ Disallow: /athletic/rss-feed/ Disallow: /athletic/*?*rss=1 Disallow: /athletic/global-color-test.php Disallow: /athletic/global-font-test.php Disallow: /athletic/graphql* User-agent: anthropic-ai Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: CCBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-SearchBot Disallow: / User-agent: Claude-User Disallow: / User-agent: Google-Extended Disallow: / User-agent: GPTBot Disallow: / User-agent: Meta-ExternalAgent User-agent: meta-externalagent Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: PerplexityBot Disallow: /
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does nytimes.com block GPTBot?
Yes. As of Oct 2, 2026, nytimes.com's robots.txt disallows GPTBot, OpenAI's training crawler.
Can ChatGPT search show results from nytimes.com?
Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.
Does nytimes.com have an llms.txt file?
We didn't find a valid llms.txt at nytimes.com/llms.txt when we checked.
Data: nytimes.com's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.