BuildWithSiddesh
Sign inSubscribe
SEO · 4 min read

Cloudflare Lets Sites Block AI Training Without Blocking Search

Learn how to use Cloudflare edge rules to block AI scrapers while keeping search engine bots. Review configuration steps for your dashboard.

· SEO leader and product builder · Updated

Cloudflare allows site owners to block artificial intelligence training scrapers while permitting traditional search engine crawlers by inspecting crawler signatures at the edge. This system relies on verified bot signatures and user agent behaviours rather than blocking entire traffic sources wholesale. Practitioners can use the AI Crawler Analyzer to review traffic patterns before applying these granular network layer blocks.

How Cloudflare separates AI crawlers from search bots

Cloudflare has introduced a feature that allows site owners to block AI training scrapers while continuing to permit traditional search engine crawlers (Cloudflare Lets Sites Block AI Training Without Blocking Googlebot). The system relies on identifying and categorising incoming requests based on verified bot signatures and user agent behaviours rather than blocking entire traffic sources wholesale.

Traditional robots.txt files often force an all or nothing choice for site administrators. A site that wanted to stop large language models from ingesting content for training purposes previously had to block search engine bots as well, which risked visibility in standard search results. Cloudflare addresses this by inspecting the crawler signatures at the edge, separating verified search bots like Googlebot from autonomous scrapers used exclusively for model training (Have it both ways: stay discoverable in search while disallowing AI training).

Practitioners managing crawl budgets and server logs can review their traffic patterns using tools like the AI Crawler Analyzer to see how these requests manifest before adjusting their edge rules. By applying granular blocks at the network layer, websites maintain organic search presence while withholding data from unauthorised machine learning pipelines.

What Cloudflare confirmed and what remains speculation

Cloudflare confirmed that site owners can now block artificial intelligence training scrapers while continuing to allow traditional search engine crawlers through a single dashboard toggle [1]. According to Cloudflare's blog, the system relies on identifying bots based on their declared intent and request patterns [2]. Verified details show that this control mechanism sits inside the Cloudflare firewall, allowing publishers to granularly separate utility traffic from model training traffic [1]. Practitioners can test how their current setup handles these requests using the AI Crawler Analyzer tool to see which bots currently access their templates.

Industry speculation remains high regarding how major search engine operators will ultimately respond to these automated classifications. While Cloudflare provides the technical means to enforce this separation, commentators note that bot operators can attempt to spoof user agents or rotate IP addresses to bypass firewall rules [4]. It is unconfirmed whether all commercial AI developers will respect these specific signals, or if they will pivot to alternative scraping infrastructures that ignore proxy level blocks [1]. Furthermore, industry reports debate whether search engines that also operate artificial intelligence models will honour the distinction between search indexing and model training when requests originate from overlapping network infrastructures [4]. Practitioners must monitor their server logs to verify whether traffic drops align with the expected bot categories or if further rule adjustments are necessary [1].

How to configure Cloudflare to block AI training today

Log in to your Cloudflare dashboard and navigate to your domain. Open the Security tab, then click on Bots. Inside the Bot Management section, locate the new rules for AI scrapers and crawlers. Cloudflare exposes specific toggles that let you separate general search engine traffic from automated systems harvesting data for machine learning models [1]. You can also review our AI Crawler Analyzer to check which bots currently hit your origin server.

To apply the restriction, select the setting designed to block AI training crawlers while explicitly permitting verified search engine bots to index your pages [2]. Cloudflare maps these controls directly to verified bot lists rather than relying solely on standard robots.txt declarations, which many aggressive scrapers routinely ignore. Once you enable the rule, the edge network intercepts requests matching known AI model signatures before they consume your server resources.

Test the configuration by checking your firewall event logs after activation. Look for blocked requests originating from known artificial intelligence user agents to confirm the edge drops them correctly while traditional search engine requests pass through to your origin. If you run custom applications or rely on specific feeds, verify that legitimate API consumers do not get caught in the broader crawler rulesets during the initial rollout.

Comparing Cloudflare controls with Google payment pilots

Cloudflare's mechanism stops AI training scrapers while keeping search traffic open, but other industry tests focus entirely on financial compensation for content usage. As reported by Search Engine Journal, Google is running an AI payment pilot that explores paying publishers for content contributing directly to AI answers. This creates a distinct split in how the market approaches content reuse.

While Cloudflare offers a defensive block paired with potential commercial frameworks like pay per crawl, Google's pilot tests a direct revenue model for inclusion in generative outputs. Publishers must evaluate whether blocking training entirely via Cloudflare's blog aligns better with their business model than waiting for prospective payment tiers from search platforms. The core difference lies in control versus monetisation. Cloudflare hands the access toggle directly to the site owner through firewall rules, whereas payment pilots keep the terms and thresholds controlled entirely by the search engine.

Practitioners assessing these models should look closely at how their analytics reflect traffic changes after implementation. If a site relies heavily on referral traffic from search engines while wanting to prevent model training, blocking scrapers via infrastructure controls preserves search visibility. Conversely, participation in payment pilots requires entering closed testing programmes where revenue terms remain opaque. For those auditing their current setup, checking server logs against tools like the AI Crawler Analyzer helps verify which bots respect these boundaries.

Frequently asked questions

Can I block artificial intelligence scrapers on my website without losing my Google rankings?

Yes, you can block artificial intelligence training scrapers while continuing to permit traditional search engine crawlers. Cloudflare introduced a dashboard toggle that relies on verified bot signatures and user agent behaviours at the edge, allowing site administrators to separate search indexing traffic from autonomous model training scrapers without resorting to an all or nothing robots.txt setup.

How does Cloudflare distinguish between standard search engine bots and AI training crawlers?

Cloudflare inspects crawler signatures at the edge by identifying incoming requests based on their declared intent, request patterns, and verified bot signatures. According to Cloudflare, this system operates inside the firewall to separate utility traffic from model training traffic, moving past standard robots.txt declarations that aggressive scrapers often ignore.

What tools can I use to check which scrapers are currently accessing my server templates?

Practitioners can review their traffic patterns using the AI Crawler Analyzer tool before adjusting any edge rules. This tool helps site owners see how specific bot requests manifest in their logs so they can verify which automated systems currently access their templates and server resources prior to activation.

Do we know if all commercial artificial intelligence developers will respect these new network blocks?

It remains unconfirmed whether all commercial artificial intelligence developers will respect these specific signals. Industry commentators note that bot operators can attempt to spoof user agents or rotate IP addresses to bypass firewall rules, and industry reports debate whether companies operating both search and AI models will honour the distinction.

Written from these sources

  1. Cloudflare Lets Sites Block AI Training Without Blocking Googlebot · searchenginejournal.com
  2. Have it both ways: stay discoverable in search while disallowing AI training · blog.cloudflare.com
  3. Google AI Payment Pilot, Search Profiles At 10,000 · searchenginejournal.com
  4. Google’s AI Payment Pilot Vs. Cloudflare & Microsoft Models · searchenginejournal.com

Put it into practice

Detect GPTBot, ClaudeBot and PerplexityBot in your server or CDN logs, verify them by reverse DNS and see what each AI crawler fetched. Free, nothing is stored.

Or put it into practice with the free SEO & AI-search tools.

Subscribe for more →