Tuesday, September 29, 2026
AI desk
/
/
Cloudflare Lets Sites Block AI Training Without Leaving Google Search

Cloudflare Lets Sites Block AI Training Without Leaving Google Search

Cloudflare’s new setting lets publishers block AI training crawlers while keeping search indexing enabled. Here is how the controls work.
Last updated
September 20, 2026
7 min read
Fact-checked

Photo: TechJournal

Share

Quick Answer

Cloudflare Disallow AI Training lets site owners reject AI model training crawlers without automatically removing their sites from search results. The September 15, 2026 control publishes a no-training preference in robots.txt while allowing accountable mixed-use crawlers to continue indexing pages for search. Site owners should review existing bot settings because older blocking options can also prevent Google, Bing, and Apple search crawling.

Key Takeaways

  • Cloudflare introduced Disallow AI Training on September 15, 2026.
  • The setting separates AI training access from ordinary search indexing.
  • Training-only crawlers from Amazon, Anthropic, Meta, and OpenAI are blocked under the new control.
  • Older Block settings can stop mixed-use crawlers, including Googlebot and Bingbot, from search crawling.
  • Cloudflare recommends different controls for Search, AI Training, and Agent traffic on ad-supported sites.

What does Cloudflare Disallow AI Training do?

Cloudflare Disallow AI Training lets a website refuse AI training use while remaining available to search crawlers. Cloudflare introduced the setting on September 15, 2026, as part of a revised set of bot controls that separates Search, Training, and Agent traffic. The distinction matters because publishers often depend on search referrals but do not want website material collected for model training.

Cloudflare says the setting publishes a no-training preference through Bot Preference Sync in robots.txt. The preference gives crawlers a machine-readable instruction about whether site content may be used for training. Accountable mixed-use crawlers can remain allowed when their operators support the preference, which means search indexing does not need to be treated as identical to AI training access. Cloudflare’s announcement describes the change and the migration path for existing settings.

The practical limitation is that the control depends on crawler operators honoring the published preference and participating in Cloudflare’s accountable approach. A robots.txt preference is not the same as a technical guarantee that every automated system will never request a page. Site owners should treat the setting as a clear policy signal and a useful enforcement control for supported crawlers, then continue monitoring bot activity through their existing security tools.

Why does separating AI training from search matter?

Separating AI training from search matters because search indexing and model training create different business and content-use consequences. A search engine indexes pages to help users find and visit the original publisher, while AI training can use content to improve a model without necessarily sending a reader back to the source website.

Cloudflare says fewer than 1% of sites on its network block search bots, while 17% use some mechanism to block training. Those figures show that many publishers want a narrower control than a universal crawler block. A broad block can reduce a site’s visibility in Google or Bing, which is a substantial tradeoff for organizations that rely on organic traffic.

Cloudflare also says mixed-use crawlers accounted for 36.6% of verified crawler traffic across its network, making them its largest crawler category. That scale explains why a single decision about a mixed-use crawler can affect both AI policy and search visibility. Publishers already evaluating how AI systems use web content may also be following the broader debate around AI audit requirements, where transparency and accountability remain central concerns.

Which AI crawlers does the new setting block?

Cloudflare Disallow AI Training blocks training-only crawlers from Amazon, Anthropic, Meta, and OpenAI without affecting search discoverability, according to Cloudflare. Training-only crawlers do not create the same search-indexing conflict because their stated purpose is collecting material for AI systems rather than returning pages in conventional search results.

Cloudflare distinguishes those training-only bots from mixed-use crawlers, which can support both search and AI-related functions. Applebot, Bingbot, and Googlebot are among the mixed-use crawlers affected differently by Cloudflare’s broader controls. Cloudflare said Apple, Google, and Microsoft either honor the setting or have committed to honor it on a specified timeline.

The important limitation is that the crawler category matters more than the brand name alone. A publisher that blocks every bot associated with a large technology company can also block a service that supports valuable search discovery. Site owners should check the stated function of each crawler and select the control that matches the site’s actual goal rather than applying the broadest available block.

Cloudflare controlPrimary purposeSearch impactPractical use
Allow SearchPermit search crawlingKeeps supported search indexing availableUse when organic discovery remains important
Disallow AI TrainingReject model training useDesigned to preserve search access for accountable mixed-use crawlersUse when a site wants search visibility without training access
BlockBlock crawler access more broadlyCan block search and training for mixed-use crawlersUse only when broad crawler blocking is intended
Block on pages with adsRestrict crawler access on ad-supported pagesCan block search for mixed-use crawlersReview carefully before applying to revenue-generating content

Can older Cloudflare blocking settings hurt Google search visibility?

Cloudflare’s older Block and Block on pages with ads settings can block search crawling as well as AI training when they apply to mixed-use crawlers. Cloudflare specifically identifies Applebot, Bingbot, and Googlebot as mixed-use crawlers under the revised framework. A site owner who previously assumed those controls only affected training traffic should review the configuration before relying on it.

The risk exists because search and training functions can be combined in one crawler category. Blocking a mixed-use crawler can therefore stop the crawler from indexing pages for search, even when the site owner only intended to reject AI model training. Search visibility may then decline after the existing indexed pages are refreshed or recrawled.

Cloudflare says existing domains using the prior Training Block or Block on pages with ads settings are migrated to Disallow AI Training to preserve the previous practical effect. The migration reduces the risk of an unexpected policy change, but site owners should still inspect their current controls. The official Cloudflare bot-control documentation explains the separate Search, Agent, and Training options.

How should ad-supported websites configure Cloudflare’s controls?

Cloudflare recommends that new ad-supported domains allow Search, disallow AI Training, and block Agent traffic on pages with ads. The recommended combination is designed to preserve search indexing while giving publishers a separate way to restrict training use and control automated agent activity around monetized pages.

Agent traffic differs from ordinary search crawling because agents may perform tasks or interact with website services rather than simply indexing a page. The distinction is relevant as more platforms add AI agents that can browse, purchase, schedule, or otherwise take actions for users. Readers tracking consumer-facing agent expansion can see the related implications of AI agents that make purchases, where automated actions create a different set of publisher and user concerns.

The recommended configuration is not a universal setting for every organization. A subscription site, an internal business portal, or a site that does not depend on search referrals may have different priorities. Website operators should first identify whether search traffic, advertising revenue, AI training restrictions, or agent access is the primary concern, then apply the matching Cloudflare control.

How can site owners check their Cloudflare AI training setting?

Site owners can check Cloudflare AI training settings by reviewing the Search, Training, and Agent controls separately before changing any existing block rule. The key task is confirming that the current setting matches the intended outcome for both training access and search discoverability.

  1. Open the Cloudflare dashboard for the relevant domain.
  2. Review the bot-management controls for Search, AI Training, and Agent traffic.
  3. Confirm whether the domain uses Disallow AI Training, Block, or Block on pages with ads.
  4. Check whether mixed-use crawlers, including Googlebot or Bingbot, remain allowed for search if search visibility is required.
  5. Review the site’s robots.txt output and bot traffic after saving a policy change.

Cloudflare’s new setting publishes a preference through Bot Preference Sync in robots.txt, so the public policy signal should align with the dashboard selection. The safest approach is to document the original configuration before changing it, particularly for a large site with advertising, international traffic, or multiple teams managing SEO and security.

Stop and consult the organization’s Cloudflare administrator or SEO lead before changing a production-wide blocking rule if the site depends on search referrals or has custom bot rules. A broad crawler block can create a visibility problem that is difficult to identify immediately, especially when different teams own revenue, editorial, and infrastructure decisions.

What does Cloudflare’s change mean for publishers and AI companies?

Cloudflare’s revised controls mean publishers have a more specific way to express consent for search indexing, AI training, and agent activity. The system addresses a practical problem: website owners have often had to choose between broad blocking and allowing crawler access without a clear distinction between the purposes behind that access.

For publishers, the benefit is more control over a site’s content policy without automatically abandoning search traffic. For AI companies, the change creates a clearer technical signal about training permission. The limitation is that the system only works as intended when crawler operators identify themselves accurately and honor the policy choices that website owners publish.

The policy change also does not resolve every dispute over web content, attribution, licensing, or AI-generated answers. Those questions involve product decisions, contracts, platform rules, and possible regulation beyond a robots.txt preference. Publishers should use Cloudflare’s controls as one part of a broader content-access policy, alongside clear terms of use, analytics monitoring, and regular reviews of how automated systems reach their pages.

FAQ

Cloudflare Disallow AI Training is designed to keep supported search indexing available while rejecting AI training use. Cloudflare says accountable mixed-use crawlers can remain allowed for search when their operators honor the no-training preference.

Which companies’ training crawlers does Cloudflare block?

Cloudflare says Disallow AI Training blocks training-only crawlers from Amazon, Anthropic, Meta, and OpenAI. The setting does not need to block search discoverability because those crawlers are categorized separately from mixed-use search bots.

Can Cloudflare’s Block setting remove a site from search results?

Cloudflare’s Block setting can prevent search crawling by mixed-use crawlers such as Googlebot, Bingbot, and Applebot. Site owners should use Disallow AI Training instead when the goal is restricting training access without broadly blocking search.

What does Bot Preference Sync do?

Cloudflare Bot Preference Sync publishes a site’s no-training preference through robots.txt. The mechanism gives supported crawlers a machine-readable policy signal about whether website content may be used for AI model training.

Should ad-supported websites block AI agents?

Cloudflare recommends that new ad-supported domains allow Search, disallow AI Training, and block Agent traffic on pages with ads. Website owners should review their own traffic and revenue needs before applying the recommendation across every page.

Share this guide
Facebook
X
LinkedIn
Written by
James Chen is a technology journalist covering artificial intelligence, software tools, and the future of work. He has been testing and reviewing AI products since 2023 and has hands-on experience with every major AI platform. His work focuses on helping everyday users get more done with AI — without the hype.

In this article

The AI Brief

Guides like this, every Friday.

One email. No hype cycle.

Keep reading