Cloudflare has announced the upcoming release of Bot Preference Sync, a new feature designed to bridge the gap between a website’s stated preferences and its enforced security rules. Available to all customers ranging from the Free tier to Enterprise, the tool automatically updates and aligns a website’s robots.txt file with the AI bot configurations established on the Cloudflare zone-level dashboard.
The launch addresses a long-standing friction point in web administration: the contradiction between what a website says it allows in its static files and how its edge security rules actually behave. Website owners frequently encounter situations where their robots.txt file disallows a specific crawler, yet enforcement rules fail to block it, or vice versa. When stated preferences and enforced rules disagree, some crawlers treat the discrepancy as justification to bypass restrictions entirely or disregard site owner guidelines.
By introducing Bot Preference Sync, Cloudflare aims to eliminate the administrative burden of maintaining multiple, separate layers of protection. Instead of requiring site owners to manually manage static files alongside edge-enforced security policies, Cloudflare will automatically generate or update the robots.txt file to reflect the specific choices made for different AI bot categories.
New Questions Facing the Internet
The launch arrives as website operators grapple with a rapidly shifting digital landscape. For years, the primary debate centered on unauthorized web scraping and whether content was being harvested to train artificial intelligence models without the explicit permission or compensation of creators. While that core issue remains pressing, website owners are increasingly forced to balance protection with discoverability and engagement.
Modern businesses must navigate complex questions about how their content appears when users pose queries to AI assistants, how to accurately track traffic coming from automated crawlers versus human visitors, and how to measure the actual economic value of referrals driven by artificial intelligence platforms.
However, these priorities vary significantly depending on a company’s underlying business model. An e-commerce merchant, for example, typically seeks maximum visibility, allowing all crawlers and training systems to index their catalog so that products surface when a consumer asks a chatbot for recommendations. Conversely, a digital publisher that monetizes content through advertisements often has entirely different incentives. Such publishers generally want to remain visible in search indexes that drive direct human readership to their articles, but they want to block their journalism from being consumed for model training—while retaining the ability to verify that their content was not utilized without authorization.
Because there is no single strategy that fits every website, Cloudflare’s tooling is designed to offer visibility and granular choice at every layer of web traffic management, ensuring that the preferences a site owner sets are accurately reflected in what they publish to the world.
The Call for Transparency

The introduction of Bot Preference Sync builds upon Cloudflare’s previous policy updates regarding mixed-use crawlers. These bots often blend search indexing, agent workflows, and AI model training behind a single user agent, putting site owners at a disadvantage by obscuring how data is collected and used.
To address this challenge, Cloudflare is tying policy enforcement directly to operational transparency. Under the updated framework, operators of bots that perform both search indexing and model training must provide specific, verifiable information about their operations to avoid being blocked when a site owner enables the "Disallow Training" setting.
Leading AI models and service providers that meet these verification criteria are tracked publicly within the AI bot transparency section of Cloudflare Radar, providing visibility into which operators honor best practices and which do not. Crawlers that fail to provide adequate transparency will not receive the benefit of the doubt and will continue to be blocked whenever a site owner disallows training.
Introducing Bot Preference Sync
Bot Preference Sync operates by reflecting the specific choices made by site owners for Search, Agent, and Training traffic on their Cloudflare dashboard directly into their robots.txt file. For websites that already utilize a robots.txt file, the preferences generated by the synchronization tool are prepended to the existing content, ensuring that any pre-existing disallow directives remain intact.
While options for Search and Agent traffic continue to include full access, blocking on ad-serving pages, or universal blocking, the setting for Training traffic provides a specialized mechanism to prevent content harvesting. When a site owner selects the Disallow option for training, a specific directive is written to the robots.txt file. Cooperating mixed-use crawlers that comply with transparency standards can still access the site for search indexing purposes, ensuring that search visibility remains unaffected while unauthorized model training is prevented.
Cloudflare utilizes data tracked within its BotBase infrastructure to periodically update the comprehensive list of bots added to the robots.txt file when a category is blocked or disallowed. Verified bots classified under Search, Agent, and Training can be reviewed at any time through the public bots directory on Cloudflare Radar.
For new customers, Bot Preference Sync will be enabled by default to streamline policy management, while existing users of the legacy managed robots.txt feature will be prompted to review their preferences and transition to the new system upon its release. Because the tool is designed to manage category-wide policies rather than complex, case-by-case exemptions, customers requiring highly customized rules can disable the synchronization feature and manually tailor their files to match specialized security arrangements.
Additionally, Cloudflare is introducing tailored onboarding defaults for ad-supported publishing sites. During initial setup, publishers who select the option indicating they monetize pages with ads will have Training automatically set to Disallow by default, making it easier to maintain search engine visibility while keeping copyrighted material out of AI training sets. Non-publisher domains will not have any blocks or disallows added automatically upon onboarding, leaving the initial configuration entirely up to the discretion of the site owner.
Bot Preference Sync will roll out to customers across all pricing tiers in the coming week.
Leave a Reply