Emerging Digital PartnersGet in touch

Article · technical

You probably blocked the AI crawlers. Most sites did it by accident.

Between Cloudflare's default settings, security plugins with a 'block AI bots' toggle, and robots.txt block lists copied from forums in 2024, a large share of business websites have opted out of AI search without anyone deciding to. Here is how it happens and how to tell if it happened to you.

By Richard Daniel3 min read

In 2024 the conversation about AI crawlers was about protecting content from being used to train models without permission. It was a reasonable concern and a lot of people acted on it: they added block lists to robots.txt, they turned on the new "block AI bots" toggles in Cloudflare and in WordPress security plugins, and they moved on.

Two years later the same crawlers, or their siblings, decide whether a business appears in ChatGPT, Perplexity and Claude when a customer asks for a recommendation. And the blocks are still there. In our audits, accidental blocking is the most common single reason a site with decent content is completely absent from AI answers. Nobody chose it. It is a default that outlived the reason for it.

Three ways it happens

The CDN did it. Cloudflare began blocking AI crawlers by default on new domains in mid-2025 and has kept adjusting the settings since. Other CDNs and hosts added similar managed rules. The result is a 403 at the edge that your robots.txt never gets a chance to override. The site looks open in every file you control and is closed in a dashboard nobody has visited.

A plugin did it. Wordfence, Sucuri, All In One Security and a dozen others shipped bot-blocking features with AI crawlers pre-listed. One checkbox, ticked once during a security review, and OAI-SearchBot has been getting a 403 ever since.

A block list did it. Someone pasted a list from a blog post into robots.txt. The list was written when GPTBot was the only OpenAI agent, so it blocks "GPTBot" and, in later versions, everything else OpenAI, Anthropic and Perplexity run. The person meant "do not train on us". The file says "do not cite us".

The distinction nobody explained

Each major AI company runs separate crawlers for separate jobs. OpenAI has GPTBot for training, OAI-SearchBot for the search index behind ChatGPT search, and ChatGPT-User for fetching a page a user has asked about. Anthropic mirrors that with ClaudeBot, Claude-SearchBot and Claude-User. Perplexity runs PerplexityBot for its index and Perplexity-User for live fetches, and says neither is used for training. Google-Extended is not even a crawler; it is a token that controls whether content fetched by Googlebot can be used for Gemini training.

Blocking a training agent is a legitimate decision about your intellectual property. Blocking a search or user agent is a decision to be absent from that engine's answers. They are different decisions, and the 2024 block lists collapsed them into one.

How to tell if it happened to you

Ten minutes:

  1. Open your robots.txt. Look for OAI-SearchBot, PerplexityBot, Claude-SearchBot and a catch-all User-agent: * with Disallow: /.
  2. Log in to your CDN and look for anything with "AI" or "bot" in the name under Security.
  3. Open your WordPress security plugin and look for a bot-blocking section.
  4. Fetch a page pretending to be a search crawler: curl -I -A "OAI-SearchBot" https://yourdomain.com/ and check for a 200, not a 403.
  5. Search your server logs for the last month for OAI-SearchBot and PerplexityBot. No hits, or only 403s, tells you what you need to know.

Make the decision on purpose

You may still decide to block training crawlers. Publishers and businesses whose content is their product often do, and OpenAI, Anthropic and Google all document how to do that without touching search visibility. What we are asking is that it be a decision. Write down your policy, set robots.txt and the CDN to match it, and check both twice a year, because the vendors keep adding agents.

The businesses winning in AI search right now are not doing anything clever. Many of them are simply the ones that were not blocked while their competitors were.

Frequently asked questions

If I block GPTBot, am I blocked from ChatGPT?

+

Not from ChatGPT search. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and live page reads use ChatGPT-User. OpenAI documents them as separate agents and says sites can allow OAI-SearchBot while disallowing GPTBot. Blocking all three is what removes you from ChatGPT answers.

Did Cloudflare really start blocking AI bots by default?

+

Yes, for new domains from mid-2025, with settings that have continued to evolve. Existing sites picked up similar behaviour through managed rule updates and one-click 'block AI bots' options. If your site is on Cloudflare and nobody has looked at the bot settings since 2024, check them.

Sources

  1. OpenAI: Overview of OpenAI crawlers · platform.openai.com
  2. Digital Heroes: Why your website isn't showing up on ChatGPT · digitalheroesco.com
  3. Anagram: AI crawler user-agent list 2026 · anagram.ai

Richard Daniel

Automation and Delivery Lead, Emerging Group

Richard leads automation and delivery across the Emerging group, working with EDP on client websites and with ETT on enterprise AI and process automation. He is the person who turns an audit finding into a working fix: crawler access, rendering, tracking, structured data and the plumbing that most marketing teams never see. He writes the technical guides on this site.

Article15 Sept 2026 · 4 min read

Your site ranks on Google and is invisible to AI. Here is why.

The most common AEO failure we find has nothing to do with content. Google renders JavaScript; the AI crawlers do not. A page can sit at position one in Search Console and be a blank shell to ChatGPT, Perplexity and Claude, and nobody notices because every dashboard is green.

Richard Daniel
technicalai-search
Article15 Sept 2026 · 4 min read

You are the only one talking about you

Every large citation study in 2026 lands in the same place: AI engines lean on Reddit, Wikipedia, YouTube, LinkedIn, editorial sites and review platforms, and cite brand-owned pages far less than owners expect. A business whose only source about itself is its own website is asking to be taken on.

George McKenna
strategyai-search
Article15 Sept 2026 · 3 min read

The schema problem: most sites have none, and the rest have it wrong

Structured data is sold as a cheat code for AI search. It is not, and the evidence says so. What it actually does is stop engines misidentifying your business and misreading your pages, and on most sites we audit it is either absent or actively wrong. Here is what we find and what to do about it.

Richard Daniel
structured-datatechnical

← All articles