Should You Block GPTBot? A Founder's Guide to AI Crawlers in robots.txt
Which AI crawlers exist, what each one does, and why most SaaS companies should allow them. Includes a robots.txt template.
· updated · 5 min read
In the last two years, dozens of sites added User-agent: GPTBot / Disallow: / to their robots.txt — often copied from a news article about publishers. For a publisher, blocking may make sense. For a SaaS company, it usually hurts.
The main AI crawlers
- GPTBot — OpenAI's crawler for model training.
- OAI-SearchBot — powers search results inside ChatGPT.
- ChatGPT-User — fetches pages when a user asks ChatGPT to browse.
- ClaudeBot — Anthropic's crawler for model training.
- Claude-SearchBot and Claude-User — fetch pages for Claude's search results and when a user asks Claude about a page.
- PerplexityBot — indexes pages for Perplexity answers.
- Google-Extended — controls use of your content for Gemini (doesn't affect Google Search).
- CCBot — Common Crawl, a dataset used by many models.
Training bots and answer bots are different
Blocking a training crawler (GPTBot, ClaudeBot, Google-Extended for Gemini training) does not, by itself, remove you from AI answers. OpenAI documents that ChatGPT search uses OAI-SearchBot, and Anthropic documents Claude-SearchBot and Claude-User for search and user requests. If you want to be cited, those answer bots matter most; whether to allow training is a separate decision.
Why SaaS should usually allow them
Your marketing site exists to be read. Buyers increasingly ask assistants for recommendations; if the assistant can't read your pages, it relies on whatever third parties say about you — or doesn't mention you at all.
A sensible template
User-agent: *Disallow: /app/Disallow: /api/User-agent: GPTBotAllow: /Sitemap: https://yoursite.com/sitemap.xml
Block your app and API, allow your marketing pages. Then verify it with our checker, and see what popular websites actually do — for example, which sites block GPTBot.