Website owners have long faced a difficult dilemma: allow proprietary content to be utilized for artificial intelligence training, or risk losing visibility and discoverability in mainstream search engine results. This friction exists primarily because some of the largest internet organizations rely on mixed-use web crawlers—single automated bots that simultaneously serve both traditional search indexing and AI model training. Under this legacy system, declining one meant rejecting both, forcing publishers and creators to choose between search traffic and intellectual property protection.
Addressing this long-standing industry tension, Cloudflare has announced a new "Disallow AI Training" setting designed to let website administrators keep their domains indexed for search while explicitly blocking those same crawlers from training on their content. Major technology firms, including Apple, Google, and Microsoft, have either honored this setting or committed to doing so within a specified timeframe.
While mixed-use crawlers represented the immediate hurdle for the training debate, the rise of AI-generated summaries presents the next major challenge for the open web. A site-wide yes or no choice remains too blunt an instrument, as the volume and context of content appearing in an AI summary matter just as much as its mere presence. Opt-outs for AI summaries are already a core requirement established for mixed-use crawler operators, and Cloudflare’s goal by early next year is to allow creators to control the precise extent of their content inclusion through a single dashboard configuration rather than negotiating with each individual operator separately.
Why Asking Isn’t Enough
The vast majority of site owners want to be found by human visitors, software agents, and constructive bots. However, a significant portion of the open internet relies heavily on advertising, subscription models, or direct visitor engagement, all of which depend on someone actually arriving at the website. While nearly every site owner views traditional search as beneficial—with fewer than one percent of Cloudflare-managed sites opting to block search bots—AI training is viewed very differently. Approximately 17 percent of sites utilize mechanisms to block training, demonstrating why administrators require granular controls instead of a generalized, one-size-fits-all approach to blocking AI.
Traditional solutions like robots.txt directives alone cannot fully resolve the problem. While anyone can publish a robots.txt file, it cannot definitively identify who is crawling a site, determine their underlying motivation, or stop a malicious crawler that simply ignores the directive. A comprehensive network infrastructure, however, can bridge this gap by publishing preferences, identifying incoming crawlers, classifying their intentions, blocking those that ignore instructions, and reporting operator behaviors publicly on Cloudflare Radar.
Simply blocking a crawler removes it from a site, but it does not change systemic crawler behavior. The preferred outcome involves operators respecting publisher boundaries without forcing a binary choice. Consequently, Cloudflare has engaged directly with crawler operators since July. The response has been encouraging, with nearly all agreeing that site owners deserve control, transparency regarding content usage, and explicit reassurances that their choices will be respected. To help publishers identify compliant partners, Cloudflare established an "Accountable" designation.
The Accountable designation recognizes both current technical capabilities and concrete commitments for future delivery. To qualify, bot operators must meet specific criteria regarding transparency and control. Industry giants Apple, Google, and Microsoft have all demonstrated that they meet these qualifications by combining existing tools with time-bound commitments for features still under development.
New Security Setting Options
Cloudflare categorizes bots according to their underlying behavior, recognizing that a single bot can exhibit multiple traits. Three distinct behaviors serve as management controls: Search, Training, and Agents.
Because mixed-use crawlers combine search and training into a single operation, they previously forced website owners into an all-or-nothing tradeoff. To eliminate this friction for Accountable mixed-use crawlers, Cloudflare introduced the Disallow AI Training setting, named after the standard Disallow directive it publishes within a site’s robots.txt file.
Previously, general "Block" and "Block on pages with ads" settings did not apply to mixed-use crawlers because blocking them could inadvertently harm search discoverability. With the introduction of the new Disallow AI Training configuration, block commands now apply universally to all training crawlers, including mixed-use variants.

Training, Search, and Agent controls are administered at the domain level. Because Disallow AI Training relies on robots.txt publishing, ad-only preferences cannot be expressed through that specific mechanism, as advertising lists are too extensive and change too frequently. Furthermore, software agents do not create the same search-discoverability tradeoff as mixed-use crawlers, and the internet lacks a universally established directive for expressing disallow preferences to agents. As open standards such as the IETF’s ai-prefs mature, Cloudflare intends to revisit agent-specific controls.
Migration and Onboarding Updates
For the vast majority of website owners, existing configurations carry over automatically without requiring manual adjustments. For legacy domains that never configured granular controls, settings migrate automatically based on previous "Block AI Bots" configurations. Sites with disabled or unselected legacy blocks transition to allowing search, training, and agents. Sites with active legacy blocks migrate to allowing search, enabling Disallow AI Training, and blocking agents on pages with ads. For domains that previously utilized granular controls, practical effects are preserved under the new definitions, transitioning previous training blocks directly to Disallow AI Training.
For newly onboarded domains, preset configurations apply depending on whether a site monetizes via advertising. Because ad revenue relies on human visitors actually viewing a page—whereas AI training substitutes a visit with an instant answer and agents fetch pages without human presence—configurations for ad-supported sites are inherently more restrictive. Sites that do not rely on ad monetization maintain open preferences for search, training, and agents, while ad-monitored sites enable preference sync, allow search, disallow AI training, and block agents on ad-serving pages. These presets remain fully customizable at any time.
Implications for Specific Mixed-use Crawlers
Major mixed-use crawlers, including Applebot, Bingbot, and Googlebot, are classified as Accountable, alongside dedicated search and training crawlers from Amazon, Anthropic, Meta, and OpenAI.
Applebot allows site owners to opt out of training by adding a Disallow rule for "Applebot-Extended" in robots.txt. Creators can also express preferences for AI summaries via the nosnippet directive in page HTML and designate paywalled content to exclude it from generative outputs. While Applebot does not yet offer URL-level inspection tools, Apple has shared details of an in-progress solution slated for next year and confirmed that disallowing training has no impact on search rankings.
Googlebot provides similar opt-out mechanisms via a Disallow rule for "Google-Extended" in robots.txt and includes a toggle within its webmaster portal to exclude content from generative search results, alongside metrics and reporting tools. Google is also developing additional URL-level transparency tools for Google-Extended and has reiterated that restricting the crawler does not affect a site’s inclusion or ranking in traditional Google Search.
Bingbot integrates granular controls and transparency via Bing Webmaster Tools, allowing site owners to express training preferences through the NOARCHIVE meta tag. Microsoft is currently building mechanisms to respect domain-level "no training" preferences directly in robots.txt, with a target release for early 2027. Until that support launches, selecting Disallow AI Training will not automatically convey a no-training preference to Bing through robots.txt, mirroring the behavior of previous training block settings. Microsoft has confirmed that using NOARCHIVE does not impact search rankings.
The Road Ahead for AI Summaries
While AI training raises foundational questions about intellectual property, model creation, and the long-term economic sustainability of original content, AI summaries introduce a separate and more immediate distribution challenge. Summaries directly impact how users discover, evaluate, and ultimately visit a business, sitting directly between a potential customer and a website by answering questions, comparing alternatives, or recommending products.
Data regarding AI summaries highlights a complex economic reality. Over half of consumers read AI summaries at the top of search results, and those users are over 40 percent more likely to end their search entirely after reading one. While this dynamic reduces overall raw visitor volume, visitors referred by AI search convert at rates ranging from three to over five times higher than those arriving via traditional search engines. AI may ultimately generate fewer visits, but it brings customers with significantly stronger intent.
As open standards like ai-prefs continue to evolve through organizations like the Internet Engineering Task Force, Cloudflare aims to provide the necessary visibility and infrastructure to help content creators navigate these shifting digital landscapes. Website owners seeking to provide feedback or engage in the conversation can reach out directly via [email protected].
Leave a Reply