Skip to content
AskableHQ
Launch

Does myanimelist.net block AI crawlers?

Blocks AI training, allows AI answers

myanimelist.net blocks at least one AI training crawler but still lets AI answer engines read its pages.

Checked from myanimelist.net/robots.txt. Robots rules change; the live file is the source of truth.

AI crawlers blocked
7 of 12
AI bots named in robots.txt
8
llms.txt
Not found
Popularity
Tranco top 5,000 (#3,196)

Can AI assistants cite myanimelist.net?

  • ChatGPT search: its search crawlers may read the site
  • Claude: its search crawlers may read the site
  • Perplexity: its search crawlers may read the site

myanimelist.net opts out of AI training by OpenAI, Anthropic, Google, Apple, Common Crawl, ByteDance, Meta. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.

AI crawler access, bot by bot

CrawlerUsed forStatus on myanimelist.net
GPTBotOpenAIAI trainingBlocked
OAI-SearchBotOpenAIAI answersSome paths restricted
ChatGPT-UserOpenAIAI answersSome paths restricted
ClaudeBotAnthropicAI trainingBlocked
Claude-SearchBotAnthropicAI answersSome paths restricted
Claude-UserAnthropicAI answersSome paths restricted
PerplexityBotPerplexityAI answersSome paths restricted
Google-ExtendedGoogleAI trainingBlocked
Applebot-ExtendedAppleAI trainingBlocked
CCBotCommon CrawlAI trainingBlocked
BytespiderByteDanceAI trainingBlocked
Meta-ExternalAgentMetaAI trainingBlocked

How myanimelist.net compares

  • 19% of the 1,278 popular sites we checked block GPTBot; myanimelist.net is one of them.
  • 18% of the 1,278 popular sites we checked block ClaudeBot; myanimelist.net is one of them.
  • 14% of the 1,278 popular sites we checked block PerplexityBot; myanimelist.net does not.
  • Among 173 sites of similar popularity, 25% block at least one AI crawler.

The robots.txt rules that apply to AI crawlers

User-agent: *
Disallow: /admin/
Disallow: /log/
Disallow: /includes/
Disallow: /comtocom.php
Disallow: /comments.php
Disallow: /ad/
Disallow: /about
Disallow: /dbchanges.php
Disallow: /clubs.php/*
Disallow: /_Incapsula_Resource
Disallow: /mymessages.php
Disallow: /sns/*/*

User-agent: AI2Bot
User-agent: Ai2Bot-Dolma
User-agent: aiHitBot
User-agent: anthropic-ai
User-agent: Applebot-Extended
User-agent: Brightbot 1.0
User-agent: CCBot
User-agent: ChatGLM-Spider
User-agent: ClaudeBot
User-agent: Claude-Web
User-agent: cohere-training-data-crawler
User-agent: Cotoyogi
User-agent: DeepSeekBot
User-agent: Diffbot
User-agent: FacebookBot
User-agent: facebookexternalhit
User-agent: Factset_spyderbot
User-agent: FirecrawlAgent
User-agent: FriendlyCrawler
User-agent: Google-Extended
User-agent: GPTBot
User-agent: ICC-Crawler
User-agent: img2dataset
User-agent: ISSCyberRiskCrawler
User-agent: Kangaroo Bot
User-agent: laion-huggingface-processor
User-agent: LAIONDownloader
User-agent: Linguee Bot
User-agent: Meta-ExternalAgent
User-agent: meta-externalagent
User-agent: meta-webindexer
User-agent: omgili
User-agent: omgilibot
User-agent: PanguBot
User-agent: SBIntuitionsBot
User-agent: Scrapy
User-agent: Sidetrade indexer bot
User-agent: Spider
User-agent: TerraCotta
User-agent: TikTokSpider
User-agent: Timpibot
User-agent: VelenPublicWebCrawler
User-agent: Webzio-Extended
User-agent: webzio-extended
User-agent: YandexAdditional
User-agent: YandexAdditionalBot
Disallow: /

User-agent: Bytespider
Disallow: /

Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.

Sites with a similar AI policy

All sites that are “blocks ai training, allows ai answers” →

Does myanimelist.net block GPTBot?

Yes. As of Oct 2, 2026, myanimelist.net's robots.txt disallows GPTBot, OpenAI's training crawler.

Can ChatGPT search show results from myanimelist.net?

Its robots.txt doesn't block OAI-SearchBot or ChatGPT-User, so ChatGPT's search crawlers may read it. Whether pages are actually shown depends on ChatGPT.

Does myanimelist.net have an llms.txt file?

We didn't find a valid llms.txt at myanimelist.net/llms.txt when we checked.

Data: myanimelist.net's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.