Cloudflare Just Redrew the Line Between AI Agents and AI Training Bots

jitendra
By
jitendra
Jitendra is a freelance writer, technical blogger, and open-source enthusiast. He closely follows emerging technologies, with a particular interest in Artificial Intelligence (AI), blockchain, and quantum...
2 Views

Starting September 15, 2026, Cloudflare will no longer treat “AI bot” as a single category. It will apply its most restrictive default yet to two specific classes of automated traffic – bots that train AI models and bots that act as agents on a person’s behalf – while leaving conventional search crawlers largely untouched. For the roughly one-fifth of the web’s traffic that Cloudflare sits in front of, that is not a minor settings change. It is a redefinition of what counts as legitimate crawling on the modern internet.

 From one switch to three

The shift traces back to July 1, 2026, when Cloudflare rolled out what it called Content Independence Day, giving every customer – including those on its free tier – the ability to sort incoming AI crawlers into three distinct categories: Search, Agent, and Training. Each maps to a different real-world use case. Search crawlers index a page so a service can answer questions about it later, and in return, publishers get referral traffic. Agent crawlers act in real time on users’ behalf, the kind of automated fetch-and-respond behavior powering a chatbot pulling live information mid-conversation. Training crawlers exist purely to harvest content that gets absorbed into a model’s weights, with no attribution and no ongoing relationship with the source site.

That three-way split matters because it replaces a blunt instrument. Until now, site owners choosing to block “AI bots” were making one binary decision that lumped together a search indexer, a live shopping agent, and a scraper feeding a training run. Cloudflare’s new taxonomy, backed by a searchable bot database called BotBase, lets publishers set independent policy for each behavior – block on every page, block only where ads run, or allow outright.

 Why September 15 is the date that counts

The July 1 changes applied only to new domains joining Cloudflare. September 15, 2026 is when the more consequential shift lands for everyone else. On that date, Training and Agent bots will be blocked by default on any page carrying advertising, across new customer signups, newly created sites from existing customers, and the entire free tier. Search bots remain permitted, on the logic that they still send visitors back to the source.

Existing paying customers get advance notice and can opt out of the new defaults through their zone security settings before the deadline. But the option to simply do nothing and keep the status quo disappears for a meaningful slice of the web’s site owners.

 The mixed-use problem nobody can dodge

The complication sits with crawlers that serve more than one purpose at once. Cloudflare’s own bot classification work found that mixed-use crawlers – bots blending search indexing with training or agent behavior in a single user-agent – now account for more than a third of all crawler activity on its network. Under the September 15 rules, any crawler exhibiting both a permitted and a restricted behavior gets judged by the most restrictive rule that applies to it.

That detail has outsized stakes for some of the web’s most recognizable bots. Googlebot, Applebot, and Bing’s crawler have historically crawled for both search indexing and AI training purposes through the same bot. Under Cloudflare’s new defaults, that dual role means all three could be blocked on ad-supported pages unless a site owner deliberately reconfigures the rule – catching search traffic that publishers still depend on in the same net meant for training scrapers. For editorial and SEO teams, that turns a policy update into an operational task: separating crawl functions into distinct, purpose-specific bots with their own user-agents becomes infrastructure work with a real deadline, not a hypothetical.

 The numbers behind the policy

Cloudflare’s rationale rests on data it has been tracking since the original Content Independence Day launch. According to the company’s own reporting, AI training now accounts for 52 percent of crawler requests on its network, up from 22 percent in the spring of 2025 – more than doubling in roughly a year. Over that same period, pure search crawling has become a shrinking share of overall bot traffic, even as it remains the crawl type publishers most want to keep.

Cloudflare frames the imbalance in blunt terms: content creators were seeing declining referral traffic even as their material fueled a growing share of AI training runs, breaking what the company describes as the three-decade-old bargain between crawlers and website owners – crawl access in exchange for referrals. Speaking at the original Content Independence Day announcement, Cloudflare co-founder and CEO Matthew Prince argued that the shift toward AI-derived content, rather than the original source, represents a structural threat to the people who produce that content in the first place.

 What comes after blocking

Cloudflare has been explicit that the categorization exercise is not really about blocking traffic outright – it is about building the plumbing for a market. The company has already opened a Pay Per Crawl marketplace and an Attribution Business Insights dashboard designed to show site owners how often a given AI model sends traffic back to their domain relative to how often it crawls that domain, arming publishers with the leverage to negotiate compensation rather than simply toggling access on or off. The classification infrastructure introduced on July 1 is what makes a category-specific pricing model technically possible, even though Cloudflare has not detailed a specific payment mechanism for Agent-category traffic yet.

 The stakes for publishers and AI companies alike

For content owners, the practical task ahead of September 15 is straightforward but not trivial: audit which bots currently touch ad-monetized pages, decide deliberately whether Search, Agent, and Training traffic should be treated differently, and opt out early if the new defaults would cut off referral-generating search access along with unwanted training and agent traffic.

For AI companies, the incentive structure flips. Crawlers that continue to blend search, agent, and training functions into a single undifferentiated bot risk losing access altogether on a large share of the ad-supported web. Separating those functions into distinct, clearly labeled crawlers – the approach Cloudflare is nudging the industry toward – becomes the price of continued access rather than an optional best practice.

The broader signal is that the era of AI companies crawling the open web on their own terms, with publishers left to negotiate from a position of near-total information asymmetry, is closing. Cloudflare is not the only infrastructure player capable of enforcing that shift, but as a company sitting in front of a meaningful share of global web traffic, its default settings now function as de facto policy for a large part of the internet – whether AI companies and publishers alike are ready for it or not.

Follow:
Jitendra is a freelance writer, technical blogger, and open-source enthusiast. He closely follows emerging technologies, with a particular interest in Artificial Intelligence (AI), blockchain, and quantum computing. Beyond writing, he loves exploring new destinations, reading books, and spending time in nature.
Leave a Comment