Does franceinfo.fr block AI crawlers?
Blocks AI training and AI answers
franceinfo.fr blocks both AI training crawlers and AI answer crawlers in its robots.txt.
Checked from franceinfo.fr/robots.txt. Robots rules change; the live file is the source of truth.
- AI crawlers blocked
- 12 of 12
- AI bots named in robots.txt
- 13
- llms.txt
- Not found
- Popularity
- Tranco top 5,000 (#2,802)
Can AI assistants cite franceinfo.fr?
- ChatGPT search: its search crawler is blocked
- Claude: its search crawler is blocked
- Perplexity: its search crawler is blocked
franceinfo.fr opts out of AI training by OpenAI, Anthropic, Google, Apple, Common Crawl, ByteDance, Meta. Training crawlers and answer crawlers are separate: blocking training doesn't by itself remove a site from AI answers.
AI crawler access, bot by bot
| Crawler | Used for | Status on franceinfo.fr |
|---|---|---|
| GPTBotOpenAI | AI training | Blocked |
| OAI-SearchBotOpenAI | AI answers | Blocked |
| ChatGPT-UserOpenAI | AI answers | Blocked |
| ClaudeBotAnthropic | AI training | Blocked |
| Claude-SearchBotAnthropic | AI answers | Blocked |
| Claude-UserAnthropic | AI answers | Blocked |
| PerplexityBotPerplexity | AI answers | Blocked |
| Google-ExtendedGoogle | AI training | Blocked |
| Applebot-ExtendedApple | AI training | Blocked |
| CCBotCommon Crawl | AI training | Blocked |
| BytespiderByteDance | AI training | Blocked |
| Meta-ExternalAgentMeta | AI training | Blocked |
How franceinfo.fr compares
- 19% of the 1,278 popular sites we checked block GPTBot; franceinfo.fr is one of them.
- 18% of the 1,278 popular sites we checked block ClaudeBot; franceinfo.fr is one of them.
- 14% of the 1,278 popular sites we checked block PerplexityBot; franceinfo.fr is one of them.
- Among 176 sites of similar popularity, 25% block at least one AI crawler.
The robots.txt rules that apply to AI crawlers
User-agent: * Disallow: /*.json Disallow: /profile/ Disallow: /mail/ Disallow: /user/ Disallow: /auth/ Disallow: /esi-block/ Disallow: /esi/www/taxonomy/ Disallow: /cgu.html Disallow: /nous-contacter.html Disallow: /charte-deontologique.html Disallow: /webview/ Disallow: /en-direct/message/ Disallow: /live/message/ Disallow: /live/iframe/ Disallow: /elections/carto/map-embed? Disallow: /histogramme-us-elector-embed*.html Disallow: /preview-examens/ Disallow: /*/recherche/?* Disallow: /*/widget_recherche.html Disallow: /*/*_resultat-eleve*.html Disallow: /*.html?callV1 Disallow: /*.html?google_editors_picks= Disallow: /*.html?*twitter_impression= Disallow: /*.html?Echobox= Disallow: /*.html?fbclid= Disallow: /*.html?gclid= Disallow: /*.html?source= Disallow: /*.html?mobile-app= Disallow: /*.html?viewType= Disallow: /*.html?title= Disallow: /*.html?print= Disallow: /*.html?*format= Disallow: /*.html?noamp= Disallow: /*.html?tblci= Disallow: /*.html?fr= Disallow: /shuffle-app/ User-agent: AI2Bot User-agent: AI2Bot-DeepResearchEval User-agent: Ai2Bot-Dolma User-agent: Amazonbot User-agent: AmazonBuyForMe User-agent: anthropic-ai User-agent: Applebot User-agent: Applebot-Extended User-agent: bigsur.ai User-agent: Bytespider User-agent: CCBot User-agent: ChatGLM User-agent: ChatGLM-Spider User-agent: ChatGPT-Operator User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User User-agent: Claude-Web User-agent: CloudVertexBot User-agent: cohere-ai User-agent: cohere-training-data-crawler User-agent: Cotoyogi User-agent: DataForSeoBot User-agent: Datenbank Crawler User-agent: Devin User-agent: Diffbot User-agent: DuckAssistBot User-agent: FacebookBot User-agent: FriendlyCrawler User-agent: gemini-deep-research User-agent: GoogleAgent-Mariner User-agent: Google-Extended User-agent: Google-NotebookLM User-agent: GPTBot User-agent: ICC-Crawler User-agent: ImagesiftBot User-agent: imageSpider User-agent: img2dataset User-agent: Kangaroo Bot User-agent: KlaviyoAIBot User-agent: laion-huggingface-processor User-agent: LCC User-agent: LinerBot User-agent: Manus-User User-agent: meta-externalagent User-agent: meta-externalfetcher User-agent: MistralAI-User User-agent: netEstate User-agent: netEstate Imprint Crawler User-agent: NovaAct User-agent: OAI-SearchBot User-agent: omgili User-agent: Omgilibot User-agent: OmigiliBot User-agent: PanguBot User-agent: PerplexityBot User-agent: Perplexity-User User-agent: PhindBot User-agent: Poggio-Citations User-agent: QualifiedBot User-agent: SBIntuitionsBot User-agent: SemrushBot-SWA User-agent: sider.ai User-agent: Spider User-agent: TavilyBot User-agent: TheKnowledgeAI User-agent: TimpiBot User-agent: TwinAgent User-agent: VelenPublicWebCrawler User-agent: webzio-extended User-agent: YouBot Disallow: /
Excerpt: only groups for all user-agents (*) or named AI crawlers, up to 40 rules per group.
Sites with a similar AI policy
Does franceinfo.fr block GPTBot?
Yes. As of Oct 2, 2026, franceinfo.fr's robots.txt disallows GPTBot, OpenAI's training crawler.
Can ChatGPT search show results from franceinfo.fr?
Its robots.txt blocks a ChatGPT search crawler, so ChatGPT search is unlikely to read or cite those pages.
Does franceinfo.fr have an llms.txt file?
We didn't find a valid llms.txt at franceinfo.fr/llms.txt when we checked.
Data: franceinfo.fr's public robots.txt and llms.txt, fetched by AskableHQ. Popularity rank from the Tranco list (ID Y83YG). Methodology.