Home / News / Cloudflare Lets Sites Block AI Training Without Killing Googlebot

GEO/AEO TrendsImpact: 70/100

Cloudflare Lets Sites Block AI Training Without Killing Googlebot

Cloudflare has introduced a 'Disallow AI Training' setting that lets site owners opt out of AI dataset harvesting without severing access for core search crawlers like Googlebot and Bingbot. The update resolves a high-stakes dilemma for business owners who want to safeguard proprietary content while preserving their rankings in organic search and AI answer engines.

VisibilityAI·11 hours ago·4 min read·Source: Search Engine Journal
Cloudflare Lets Sites Block AI Training Without Killing Googlebot

Key Highlights

  • Cloudflare introduced a 'Disallow AI Training' setting that keeps search indexation open while restricting LLM training.
  • Prevents websites from inadvertently blocking Googlebot, Bingbot, and Applebot when attempting to protect content.
  • Replaces legacy binary toggles with three granular controls: Search, Training, and Agent.
  • Users on previous blocking rules are automatically migrated to safe settings to protect existing SEO traffic.

Until recently, website owners faced a frustrating catch-22: protect original content from being scraped by AI developers, or risk crippling organic search visibility. Because major tech platforms rely on the same crawlers—specifically Googlebot, Bingbot, and Applebot—for both search indexation and AI model training, blocking AI bots previously meant potentially locking out core search engines altogether.

Web infrastructure provider Cloudflare has stepped in with a practical solution. Its new Disallow AI Training control communicates a strict 'no-training' preference while leaving your site fully accessible to search crawlers.

Here is a breakdown of what changed, how Cloudflare handles mixed-use bots, and what your brand should consider when balancing content defense against Generative Engine Optimization (GEO).

What Happened

Earlier this summer, Cloudflare warned that stricter bot-blocking settings would soon sever connections with mixed-use crawlers. Under an earlier proposal slated for mid-September, choosing to block AI training was set to lock out Googlebot, Bingbot, and Applebot completely. That blunt approach reflected how tech giants often bundle search indexing and data collection for large language model (LLM) training under the same user-agent.

Recognizing that cutting off Googlebot would devastate commercial websites, Cloudflare revised its framework. Rather than relying on an all-or-nothing switch, the company introduced Disallow AI Training.

This setting appends an explicit no-training directive to your robots.txt file. The directive informs crawlers that content can be indexed for organic search results, but must not be harvested for machine learning training datasets. Crawlers classified as 'Accountable' by Cloudflare—meaning they honor webmaster directives—are permitted to continue crawling for organic search visibility.

Key Details: The New Crawler Controls

Cloudflare is replacing its legacy 'Block AI Bots' toggle and Managed Robots.txt features with three distinct bot control categories:

  • Search: Determines whether standard search engines can index your site for typical SERP results.
  • Training: Governs whether AI operators can scrape your pages to train foundation models (like GPT-4, Gemini, or Claude).
  • Agent: Manages autonomous AI agents executing actions or retrieving live data on behalf of end users.

How Existing Sites Are Handled

To keep businesses from unintentionally vanishing from Google Search, Cloudflare is automatically reconfiguring existing accounts:

  • Sites that previously chose to 'Block' or 'Block on pages with ads' under the training setting are being automatically shifted to Disallow AI Training.
  • Accounts using the older 'Block AI Bots' toggle will now default to Allow for Search, Disallow AI Training for Training, and Block on pages with ads for Agent.
  • If a site administrator manually insists on setting the Training toggle to full Block, Cloudflare will completely block Googlebot, Bingbot, and Applebot, which will eliminate your rankings in search results.

The GEO Trade-Off: Training Data vs. Real-Time Citations

For businesses focused on Generative Engine Optimization (GEO)—earning citations and recommendations in tools like ChatGPT, Perplexity, and Google AI Overviews—this update introduces an important strategic question: Should you disallow AI training?

1. AI Overviews & Live Search (RAG): When AI platforms generate direct answers using Retrieval-Augmented Generation (RAG), they pull directly from live search indices via Googlebot, Bingbot, or dedicated search agents. By permitting search while disallowing training, you retain eligibility to appear in AI Overviews and Perplexity search answers.

2. Base Model Knowledge: Restricting AI training prevents your proprietary research, case studies, and whitepapers from being baked into the underlying weights of future foundation models. While this protects your copyrighted work, it also means offline or zero-search AI prompts will lack direct knowledge of your business.

What It Means For Your Business

For most local and mid-sized businesses, a website exists primarily to drive customer acquisition and inbound visibility rather than to license raw training data.

  • Check Your Cloudflare Dashboard: Verify your current settings under Cloudflare's Bot Management tab. Ensure your Search control is set to Allow so your organic visibility remains intact.
  • Evaluate Your Content Strategy: If you produce highly proprietary research, code, or creative media, using Disallow AI Training allows you to guard your assets against foundational LLM scrapers without sacrificing traditional SEO.
  • Don't Cut Off Your Discovery Channels: If your priority is brand discovery in the AI era, avoid setting mixed-use crawlers to a hard 'Block'. Complete blocks prevent both Google Search rankings and real-time AI citation engines from recommending your services.

Why This Matters For Your Business

This update eliminates one of the sharpest operational headaches facing site operators: the threat of tanking hard-earned search traffic simply by trying to defend proprietary content. Most businesses rely heavily on Googlebot and Bingbot for inbound customer discovery. When Cloudflare previously warned that blocking AI training might require locking out mixed-use crawlers entirely, business owners faced an impossible choice between content scraping and commercial invisibility. The shift also brings stability to Generative Engine Optimization (GEO). Systems like Google AI Overviews and Bing Copilot depend on standard search pipelines to fetch live sources for generated summaries. Establishing a clear separation between search indexing and model ingestion allows companies to remain discoverable in real-time retrieval tools without passively surrendering proprietary insights to corporate training datasets. Ultimately, this gives growth-focused businesses clarity. You can now safeguard your intellectual property without putting your primary organic acquisition channels on the line.

Frequently Asked Questions

Will enabling 'Disallow AI Training' hurt my Google search rankings?

No. Under Cloudflare's updated framework, Googlebot is classified as an accountable mixed-use crawler. It can still index your site for organic search results while honoring your preference not to have content harvested for base model training.

Can I still appear in Google AI Overviews if I disallow training?

Yes. Google AI Overviews rely primarily on Google's search index and live retrieval to answer user queries. As long as Googlebot has 'Search' permissions enabled, your content remains eligible for AI search snippets and traditional rankings.

What happens if I set the Training control to full 'Block'?

If you choose a hard 'Block' instead of 'Disallow AI Training', Cloudflare will sever access for mixed-use crawlers entirely. Locking out Googlebot, Bingbot, and Applebot in this manner will remove your website from organic search engine results and real-time AI citations.

Is your business showing up in AI search?

Get your free AI visibility audit - see if ChatGPT, Perplexity, and Google AI actually recommend you.