Emerging Digital PartnersGet in touch

How-to guide · technical

How to set up robots.txt for AI crawlers without locking yourself out of AI answers

A working robots.txt template for 2026 that keeps you visible in ChatGPT, Perplexity, Claude, Gemini and Copilot, lets you make a separate decision about model training, and avoids the copy-and-paste block lists that quietly remove sites from AI search.

By Richard Daniel4 min read

Your robots.txt decides whether AI engines are allowed to fetch your pages. Get it wrong in one direction and you are training other people's models for free with no visibility in return. Get it wrong in the other and ChatGPT, Perplexity and Claude cannot cite you at all. Most sites we audit have not made a decision either way; they have inherited a file from a theme, a plugin or a block list someone pasted in 2024.

This guide gives you a decision framework, the agent names to use (spelled exactly as the vendors publish them), and a template you can adapt.

Step 1: Understand the three jobs an AI bot can have

Every major AI company runs more than one crawler, and they do different things.

CompanyTraining (builds the model)Search / index (powers cited answers)User fetch (loads a page live for a user)
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexitynone declaredPerplexityBotPerplexity-User
GoogleGoogle-Extended (opt-out token, not a separate bot)Googlebot (AI Overviews and AI Mode use the normal index)n/a
Microsoftn/aBingbot (ChatGPT search draws heavily on Bing's index)n/a
AppleApplebot-Extended (opt-out token)Applebotn/a
MetaMeta-ExternalAgentn/an/a
AmazonAmazonbotAmazonbot (Alexa answers)n/a
Common CrawlCCBotn/an/a
ByteDanceBytespidern/an/a

Blocking a training agent is a business decision about your content being used to build models. Blocking a search or user agent is a visibility decision: you will not appear in that engine's answers. Treat them separately.

Step 2: Decide your policy

Three sensible positions:

  1. Open: allow everything. Simplest, maximum visibility, your content may be used for training. Right for most small businesses whose content is marketing rather than a product.
  2. Visible but not training: allow all search and user agents, disallow the training agents. Right for publishers, consultancies and anyone whose content is the product.
  3. Closed: block all AI agents. Only right if you have decided AI visibility does not matter to you. Be honest about whether that is a decision or a default.

Whatever you pick, Googlebot and Bingbot must stay allowed. They power the web index behind Google's AI features and ChatGPT search respectively.

Step 3: Write the rules

The template below is position 2. For position 1, delete the Disallow: / lines under the training agents. Each User-agent group needs at least one Allow or Disallow line or it does nothing.

# --- Search and retrieval crawlers: keep these allowed to appear in AI answers ---
User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

User-agent: Applebot
Allow: /

User-agent: Amazonbot
Allow: /

# --- Training crawlers: this section opts out of model training only ---
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Bytespider
Disallow: /

# --- Everyone else, including Googlebot and Bingbot ---
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/

Sitemap: https://yourdomain.com/sitemap.xml

Two details matter. Spelling must match the vendor's token exactly (GPTBot, not GPT-Bot). And a bot only uses the most specific group that names it; if you have a User-agent: * group that disallows a folder, a named agent with its own group does not inherit that rule, so repeat any path exclusions you care about inside the named groups.

Step 4: Make sure the firewall agrees

Robots.txt is read after the request reaches your server. If Cloudflare, your host or a security plugin returns a 403 first, the file is irrelevant. After editing robots.txt, check the bot settings on your CDN and any WordPress security plugin, and switch off blanket "block AI bots" toggles if your policy is to be visible. Then confirm in your logs that OAI-SearchBot, PerplexityBot and Claude-SearchBot get 200s.

Step 5: Test it

  • Google Search Console has a robots.txt report under Settings that shows the fetched file and any parse errors.
  • Bing Webmaster Tools has a robots.txt tester.
  • For AI agents, fetch the file yourself and check the groups are separated by blank lines with no stray characters. Then run the curl test from our guide on checking whether AI engines can read your site, using the OAI-SearchBot user agent, and confirm a 200.

Step 6: Review it twice a year

Vendors add agents. OpenAI split GPTBot and OAI-SearchBot; Anthropic added Claude-SearchBot; Perplexity added Perplexity-User. Put a calendar reminder in for January and July to check the vendor docs linked below and update the file.

Common mistakes

  • Pasting a "block all AI" list from a forum in 2024 that includes OAI-SearchBot and PerplexityBot, then wondering why the business never appears in ChatGPT.
  • Blocking Google-Extended believing it protects rankings. It has no effect on Search; it only limits Gemini training use.
  • A User-agent: * group with Disallow: / left over from a staging site. Every AI bot without a named group obeys it.
  • Assuming robots.txt stops bad actors. It does not. Use rate limiting and WAF rules for that.

Frequently asked questions

Can I block AI training but still appear in AI search?

+

Yes. Disallow the training agents (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider) and allow the search and user agents (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User) plus Googlebot and Bingbot. OpenAI, Anthropic and Perplexity all document these as separate agents.

Do all AI bots respect robots.txt?

+

No. The major vendors say their search and training crawlers do, but user-triggered fetchers such as ChatGPT-User and Perplexity-User state that robots.txt may not apply because a human requested the page, and some third-party scrapers ignore it entirely. Robots.txt is a request, not a lock. Use firewall rules for bots that misbehave.

Should I block Google-Extended?

+

Google-Extended only controls whether your content can be used for Gemini training and grounding. It does not affect Google Search, AI Overviews or AI Mode, which use Googlebot. Blocking Googlebot would remove you from all of them, so never do that.

Sources

  1. OpenAI: Overview of OpenAI crawlers · platform.openai.com
  2. Anthropic: Does Anthropic crawl data from the web? · support.anthropic.com
  3. Anagram: AI crawlers explained and how to let them in · anagram.ai

Richard Daniel

Automation and Delivery Lead, Emerging Group

Richard leads automation and delivery across the Emerging group, working with EDP on client websites and with ETT on enterprise AI and process automation. He is the person who turns an audit finding into a working fix: crawler access, rendering, tracking, structured data and the plumbing that most marketing teams never see. He writes the technical guides on this site.

Guide15 Sept 2026 · 3 min read

How to get indexed by Bing (because ChatGPT search leans on it)

OpenAI names Bing among the search providers behind ChatGPT search, and independent analyses in 2026 found most of ChatGPT's cited pages also rank in Bing's top results. Copilot is Bing. If your Bing coverage is thin, so is your ChatGPT visibility. Here is the fix.

Richard Daniel
technicalai-searchmeasurement
Guide15 Sept 2026 · 3 min read

How to write an llms.txt file (and what it does and does not do)

Google says you can ignore llms.txt. Ahrefs found 97 percent of the files it studied were never requested. No major AI lab has committed to reading it. It is still a 30-minute job with no downside, as long as you know exactly what you are and are not getting. Here is how to do it properly.

Richard Daniel
technicalai-search

← All guides