SEO/GEO

Should You Block AI Bots (GPTBot, ClaudeBot, PerplexityBot) on Your Site in 2026?

By 18 August 2026No Comments

Key takeaways:

  • Blocking AI bots is not a single decision: three families of crawlers exist (training, search, user agent) and four blocking levers that do not control the same things.
  • On September 15, 2026, Cloudflare will block Googlebot, Applebot and Bingbot for every customer who has enabled training blocks, including through the legacy “Block AI bots” switch from July 2025.
  • Google-Extended does not remove you from Google Search or from AI Overviews: it only controls training for Gemini and Vertex AI.
  • For a small or mid-sized business, the risk is not being copied by an AI, it is not being cited at all.

No, unless your content is your product. Blocking AI bots is decided bot by bot, and four distinct levers exist. Three of them changed between July and September 2026.

Measured by Cockpyt AI
Does ChatGPT recommend your brand?Measure your presence and identify the brands cited in your place. No credit card required.

Start a 14-day free trial →

On July 1, 2026, Cloudflare replaced its single “Block AI bots” switch with three categories: Search, Agent and Training. On September 15, blocking the Training category will also block Googlebot, Applebot and Bingbot, which crawl the web for both purposes at once (Cloudflare, “Your site, your rules,” July 1, 2026). A box ticked out of caution a year ago could therefore push you out of Google within weeks. The robots.txt file, GPTBot and ClaudeBot are only part of the problem.

Should you block AI bots? The one-line answer

No for the vast majority of business websites, yes in one case only: when your content is the product you sell.

The question is framed wrongly as soon as you speak about “AI bots” in the singular. There is no homogeneous category of crawlers to allow or refuse as a block. There are three distinct behaviours, carried by different crawlers, with opposite consequences for your business.

A brochure site, a consulting firm, an online shop or a tradesperson lives on AI visibility. Their problem is not that a model learns their pricing page. Their problem is that ChatGPT cites three competitors and not them when a prospect asks for a recommendation. Blocking citation crawlers means voluntarily withdrawing from the commercial conversation.

A news publisher, a paid database, a training platform whose courses make up the sellable catalogue sit in the opposite position. Their content has direct market value, and a model that reproduces it for free destroys that value. Blocking becomes an asset protection decision, not a technical setting.

Between those two, the question is settled line by line, bot by bot, with the right lever.

The 4 blocking levers and what each one actually controls

Four mechanisms coexist in 2026, and they differ in both scope and force. Confusing them produces most of the errors I encounter during audits.

Lever Nature What it controls Scope
robots.txt Declarative Which user agents may crawl which URLs Every crawler that chooses to respect it
WAF / CDN (Cloudflare, Fastly, application firewall) Enforced Actual server access: the request receives a 403 response Every crawler, including those ignoring robots.txt
Google-Extended Declarative Downstream use: training and grounding for Gemini and Vertex AI Google only, no effect on Google Search
Search Console “Search generative AI” Enforced on Google’s side Your presence in AI Overviews, AI Mode and Discover Google only, no effect on blue links

The distinction that matters: robots.txt is a request, the firewall is a decision. When the two contradict each other, the firewall wins, and your robots.txt becomes a document describing a policy that is no longer in force. I have seen sites whose robots.txt explicitly allowed PerplexityBot while their host returned a 403 to it for months.

Major operators state that they respect RFC 9309, the standard governing robots.txt. Less established players ignore it. Real blocking therefore always runs through the server layer, never through a text file.

Which bots do what? The 2026 map

Three families of behaviour, to be handled separately. Cloudflare built its new classification on exactly this distinction: “Search,” “Agent” and “Training.”

Training crawlers collect your content to train or fine-tune a model. They send back no traffic and produce no citation.

  • GPTBot (OpenAI)
  • ClaudeBot (Anthropic)
  • CCBot (Common Crawl, whose archives feed many models)
  • Google-Extended and Applebot-Extended: these are control tokens, not crawlers. You will never see them in your logs.
  • Meta-ExternalAgent (Meta)

Search and citation crawlers build an index to answer user questions, with a link back to your site. These are the ones producing your AI referral traffic.

  • OAI-SearchBot (ChatGPT Search)
  • PerplexityBot (Perplexity)
  • Claude-SearchBot (Anthropic)

User agents fetch a page in real time because a human has just asked a question. Someone is waiting behind the request.

  • ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User

Blocking a user agent produces a particularly counterproductive effect: a prospect explicitly asked to see your page, and you refuse them access. On an online shop, this family of crawlers is starting to complete full purchase journeys.

September 15, 2026: what Cloudflare changes without telling you

On July 1, 2026, Cloudflare retired its single AI bot blocking switch, replacing it with three independent settings. The timeline deserves your immediate attention.

For new domains onboarding to Cloudflare, the Training and Agent categories will be blocked by default on pages displaying ads, while the Search category remains allowed. The stated logic is simple: an ad signals that a human was expected on that page.

The genuinely dangerous point lies elsewhere. From September 15, multi-purpose crawlers will be evaluated according to all of their behaviours, and the most restrictive rule applies. Cloudflare states it without ambiguity: multi-purpose crawlers such as Googlebot, Applebot and Bingbot will be blocked for customers who have selected to block training, whether through the new options or through the legacy “Block AI bots” service.

The trap to check this week

If you enabled “Block AI bots” on Cloudflare in July 2025 to protect yourself from ChatGPT, that setting will block Googlebot from September 15, 2026. Opting out is done in your Cloudflare zone security settings, before that date.

Cloudflare is also testing a use signal extending Content Signals in robots.txt, with three values: immediate (store nothing), reference (index, excerpt and link back, the default value) and full (summarise and reproduce). This signal expresses a preference, it blocks nothing by itself.

In my view, the real risk for a business has never been deliberate, carefully considered blocking. It is inherited blocking: a box ticked by a contractor eighteen months ago, in a media climate where “blocking AI” looked obvious, and that nobody has read since. These settings sleep in Cloudflare accounts the owner never opens. September 15 will expose those oversights in the worst possible way, through a drop in organic traffic that takes weeks to trace back to its cause.

Google-Extended does not remove you from AI Overviews

Google-Extended controls one thing only: the use of your content to train and ground Gemini and the Vertex AI generative APIs. Google states that this token affects neither inclusion in Google Search nor ranking, and that it is not used as a ranking signal.

No crawler named “Google-Extended” visits your site. Googlebot does the crawling, as before; the token only indicates what Google may subsequently do with the retrieved content. Blocking it therefore does not reduce your server load and changes nothing about your indexing.

The consequence surprises many business owners: blocking Google-Extended does not remove you from AI Overviews or AI Mode. Those surfaces are part of Google Search and draw on the standard index. To appear in them, a page only needs to be indexed and eligible for a snippet.

Google launched AI Overviews and AI Mode in France on July 22, 2026, after more than two years of delay tied to the neighbouring rights dispute (Google France, July 22, 2026). A dedicated control now exists in Search Console, under the “Search generative AI” label. It lets you choose whether your site feeds those surfaces. Google specifies that this setting is not used as a ranking signal for the rest of search.

The trade-off is clear: a site that opts out receives neither traffic nor impressions from AI Overviews and AI Mode, while keeping its position in the blue links. For a commercial site, this means voluntarily leaving a fast-growing discovery channel. For a publisher whose articles are fully summarised above the results, the calculation can flip.

Block or allow? The decision tree by profile

Four profiles, four different answers. Identify yours before touching a file.

Brochure site, services, freelancers and small businesses

Allow everything. Your content has no standalone market value: it demonstrates competence and generates quote requests. Every citation in ChatGPT, Perplexity or Gemini acts as a recommendation in front of a prospect at decision stage. The server cost of these crawlers stays marginal on a site of a few hundred pages.

E-commerce

Allow search and agents, arbitrate training. User agents consult your product pages on behalf of a buyer who is comparing options. Blocking them means shutting the door on a customer who is knocking. Training is worth discussing if your catalogue contains high-value descriptions written in house.

Publishers and ad-monetised sites

Block training, allow search, arbitrate agents. Your model rests on page views, and a summary answering in place of your article deprives you of the associated revenue. Check the Googlebot trap described above before enabling training blocks at your CDN.

Proprietary knowledge base or paid content

robots.txt is not enough and never has been. Content that must not circulate is protected by authentication, access control and enforceable terms of use. A text file at the root of your domain remains a convention between well-behaved parties.

In my view, the debate about model training is a false debate for the overwhelming majority of small and mid-sized businesses. Your 40 brochure pages weigh nothing in a training corpus, and their absence will change nothing about the model’s capabilities. Your absence from generated answers, on the other hand, changes everything about your revenue. The real subject in 2026 is not “how do I stop AI from reading me,” it is “how do I make sure it cites me correctly.” I have seen far more companies lose business by blocking themselves in error than by being too open.

The configuration I recommend in 2026

Here is the baseline I deploy for clients who are not paid content publishers. Adapt it, do not copy it without thinking.

# Search and citation crawlers: allowed
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

# User agents: allowed
User-agent: ChatGPT-User
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Perplexity-User
Allow: /

# Training crawlers: arbitrate based on your model
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: CCBot
Disallow: /

# Traditional search engines
User-agent: Googlebot
Allow: /

User-agent: Bingbot
Allow: /

Sitemap: https://your-domain.com/sitemap_index.xml

Three notes on this file. An explicit Allow: / is strictly unnecessary under the standard; it serves as documentation for the next developer who opens the file. CCBot is the only crawler I block by default, because Common Crawl redistributes its archives with no control over final use. Exclusion rules for your admin pages, faceted navigation and URL parameters remain necessary on top of these.

On the Cloudflare side, the matching configuration means allowing Search and Agent, then deciding separately about Training. If you choose to block Training, go to your zone security settings and explicitly flag that you want no changes to multi-purpose crawlers, otherwise Googlebot will fall with them.

How to check that your decision is actually applied

An unverified blocking policy is worth nothing. Four checks, in this order.

  1. Test real access with a declared user agent. From a terminal, simulate a crawler visit and observe the HTTP response code. A 200 confirms access, a 403 signals a firewall block.
  2. Compare your robots.txt against your server logs. Look for the user agents you believe you are allowing. Their complete absence from the logs over thirty days indicates upstream blocking, at CDN or host level, regardless of what your file says.
  3. Open your CDN dashboard. On Cloudflare, the AI traffic management page shows the state of all three categories. That page is authoritative, not your robots.txt.
  4. Measure the outcome, not the setting. A correctly configured block shows up in citations: your brand gradually disappears from generated answers. A correctly configured allowance reads the same way, in reverse.

The command to run for the first check:

curl -A "OAI-SearchBot" -I https://your-domain.com/

The fourth check is the only one linking a technical decision to a commercial effect. A correct setting in an interface proves nothing until you observe its effect on your citations.

Measured by Cockpyt AI
Does ChatGPT recommend your brand?Measure your presence and identify the brands cited in your place. No credit card required.

Start a 14-day free trial →

Disclosure: I am a co-founder of Cockpyt AI, a French tool tracking brand visibility in ChatGPT, Perplexity and Gemini. Tracking for Google AI Overview and AI Mode is in development, with no release date announced.

Frequently asked questions

Does blocking GPTBot lower my Google rankings?

No. GPTBot belongs to OpenAI and has no connection to Googlebot or to Google’s index. The only mechanism capable of dropping your Google rankings while you think you are blocking an AI is the Cloudflare setting described above, which from September 15, 2026 takes Googlebot down along with the Training category.

Do AI bots really respect robots.txt?

Major operators state that they respect RFC 9309 and publicly document their user agents. Less established crawlers ignore it, and some present a misleading identity. A robots.txt expresses a preference; only a rule at the application firewall level produces effective blocking.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects data to train OpenAI’s models and returns no traffic. OAI-SearchBot indexes your pages for ChatGPT Search and lets your site be cited with a clickable link. Blocking the first while allowing the second is the most common configuration among publishers.

How do I know whether my host is already blocking AI bots without my knowledge?

Run a curl request with an AI crawler user agent and check the HTTP code. A 403 or a challenge page confirms server-side blocking. Then check your CDN dashboard: many accounts inherited a default block enabled during earlier waves, with no action from the site owner.

Does the llms.txt file block AI?

No. The llms.txt file is a proposed format meant to guide models toward your main content, not an exclusion mechanism. It carries no normative weight and no major engine has announced treating it as a blocking directive. Access control runs through robots.txt and the firewall.

Does blocking AI bots legally protect my content?

A technical block documents your objection to the use of your content, which can serve as evidence, without constituting legal protection in itself. Your copyright exists independently of your robots.txt. For high-value content, combine explicit terms of use, access control and server-side blocking. This article does not replace advice from a specialist lawyer.

Sources

Facts verified on August 11, 2026. AI crawler access rules evolve quickly: the new Cloudflare settings take effect on September 15, 2026.

  • Cloudflare, “Your site, your rules: new AI traffic options for all customers,” July 1, 2026 : blog.cloudflare.com
  • Cloudflare Developers, “New options to manage AI traffic,” changelog of July 1, 2026 : developers.cloudflare.com
  • Google France, Sébastien Missoffe, “Launch of AI Overviews and AI Mode in Google Search in France,” July 22, 2026 : blog.google
  • Google Search Central, Google-Extended and generative engine crawler documentation : developers.google.com
Florian Zorgnotti

As a WordPress SEO Consultant in Nice and co-founder of Cockpyt AI, I support infopreneurs, small businesses, and SMEs in their web marketing strategy and their search for online visibility. Specialized in WordPress SEO, I also offer coaching and online training. My LinkedIn profil

Leave a Reply

Favicon
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.