Cloudflare changes what its AI blocking controls do on September 15, and the change most likely to catch site owners off guard is not the new default setting. The one-click preset to block AI bots, described by the company as covering single-purpose training crawlers, will start catching Googlebot.

From that date, according to Cloudflare’s announcement, crawlers that do more than one job will be “allowed/blocked according to all of their behaviors,” and the defaults “will be enforced by the most restrictive applicable rules.” That brings in multi-purpose crawlers such as Googlebot, Applebot and BingBot, which “will be blocked by customers who have selected to block Training.” That applies whether the block was set through the new controls or through the older Block AI bots service. A site that turned that preset on to keep its pages out of model training will begin turning away the crawler behind Google Search without anyone touching a setting.

Key facts

  • The new default: on pages that display ads, training and agent crawlers are blocked by default from September 15, while search crawlers stay allowed.
  • Who gets the new defaults: new customers, new sites on existing accounts, and every existing free-plan customer that has not changed its settings by that date.
  • How far the Google block reaches: it follows whatever training block a site already set, which can cover all pages or only those that carry ads.
  • The way out: owners can decline the new defaults in their Security settings at any point before September 15.

Three Categories, and a Default Tied to Ads

Cloudflare has split automated traffic into three categories that describe what a crawler does rather than who operates it. Search is “any behavior that collects or indexes your content, so it can answer questions about it later.” Agent is “automated behavior that is acting, usually in real time, on a person’s behalf, to get something done right now.” Training is “a crawler taking your content to train or fine-tune a model.”

On September 15 the defaults attached to those categories change. In Cloudflare’s words, “the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.” The company ties the rule to advertising deliberately, arguing that “an ad is a signal that a website owner meant for a person to land there and see it.”

The reach is wider than the phrase “new defaults” suggests. Cloudflare’s announcement describes them as applying to domains newly onboarding to the service. Its press release goes further and sets out three groups: new customers, new sites added by existing ones, and, on the same date, all existing free-plan customers that have not changed their settings. Owners who would rather not receive them can say so in their Security settings at any time before September 15.

A quieter change has already taken effect. When Cloudflare announced the new controls on July 1, it also changed what the verified label buys a bot. Previously all such bots were allowed by default. The company says it is now “no longer viewing Verified as ‘default allowed’,” so what a crawler may do follows from its category rather than from its badge.

The Preset That Now Reaches Google

The older preset reaches further now because the classification behind it changed. Cloudflare describes the managed option it shipped in 2025 as one that “included single-purpose bots that crawled data for model training.” The change now, it says, recognizes that bots with multiple purposes “should be tracked with all purposes, not just one of them.” From September 15 the most restrictive rule that applies to any one of a crawler’s behaviors governs the whole crawler.

How far that reaches depends on the setting already in place. Cloudflare lets each preset apply to every page, only to pages that display ads, or to nothing at all. A training block set site-wide turns Googlebot away site-wide, while one limited to pages carrying ads stops it only there.

Google sits at the center of that because of how it crawls. In its own report, Cloudflare puts Google at approximately 88% of referral traffic. It adds that while most leading AI companies separate discovery crawlers from training crawlers, Google does not, and argues this leaves it with “about 2x more information than leading AI companies.”

Google has pushed back against that characterization before, as TechCrunch reported, pointing to Google-Extended. It is not a separate crawler. Google’s documentation describes it as a robots.txt token, with the crawling still done by existing Google user agents. It lets publishers control whether the content Google collects may be used to train future Gemini models, and for grounding in Gemini Apps and Vertex AI. Using it does not affect a site’s inclusion in Google Search.

Cloudflare sorts a crawler by everything it does. Google’s position is that owners already hold a separate control over part of what happens to the content afterward. Googlebot itself crawls for Search including AI features such as AI Overviews and AI Mode.

Cloudflare is explicit about what it wants the change to achieve. “We hope that our proposed default changes encourage mixed-use crawlers to separate out search from agent use and training,” chief executive Matthew Prince told TechCrunch.

What a Provider Decides Before September 15

Cloudflare puts its own footprint at more than 20% of the web, and says 36% of the world’s most-visited websites rely on its network. A change to its defaults is therefore not a narrow event. Among existing customers, the new defaults arrive unprompted only on the free plan. The rule on multi-purpose crawlers is not limited by plan, and applies wherever a training block has been selected.

For a provider that administers zones on behalf of customers, the deadline sets up a choice rather than a task. For zones the new defaults cover, leaving the settings alone lets them take effect. Separately, any training block already in place begins reaching multi-purpose crawlers such as Googlebot. Marking the opt-out before September 15 keeps the current behavior. Cloudflare places that control in each zone’s Security settings.

What the choice turns on can be established in advance: which zones carry a training block today, whether that block covers the whole site or only pages that display ads, and whether the customer set it themselves. An owner who switched it on to keep their pages out of model training did not necessarily mean it to reach search.

The Numbers Cloudflare Used to Make the Case

The measurements published alongside the announcement describe a crawl mix that no longer matches the labels the old controls were built around.

MeasureCloudflare’s reported figure
Crawler requests made for AI training, June 202652%, up from 22% in spring 2025
Share of crawler activity from mixed-use crawlersOver 36%
Non-human share of internet trafficOver 50%, crossed for the first time in 2026
Human traffic to some of the most heavily crawled categoriesDown as much as 40% in under a year
Google’s share of referral trafficApproximately 88%

These figures describe different cuts of crawler activity rather than shares of a single total. The first attributes crawler requests by purpose. The second measures the share of activity coming from mixed-use crawlers. They are not meant to be added together.

Every one of them is the company’s own, compiled from Cloudflare Radar and its Investor Day 2026 presentation. That matters for how much weight they carry. Cloudflare is measuring a problem it also sells the remedy for, and we have no independent source against which to check the figures.

Pay Per Crawl Is Still in Closed Beta

Cloudflare launched the one-click block option and a pay-per-crawl marketplace together in July 2025. Fourteen months on, the enforcement half is available to every customer and the monetization half is not. Pay Per Crawl answers an unpaid request with a price and the status code HTTP 402 Payment Required. Cloudflare’s own documentation still describes it, as of September 11, as being in closed beta.

The two do not combine freely. Cloudflare’s documentation states that where a crawler is blocked through its firewall or bot management products, those rules override the charge behavior. For a given crawler the blocking rule wins, and the same request cannot be both blocked and billed.

What exists in between is a licensing market that Cloudflare counts at more than 50 publisher-AI agreements signed since 2023. Licensing today, the company writes, “remains largely bespoke and unlikely to fully replace lost referral, advertising, and affiliate revenue.” Site owners have a finer set of controls than they had a year ago. What replaces the traffic those controls are a response to is still unsettled.

About the Data

The three categories, the new defaults, the rule for crawlers that do more than one job, the opt-out and the July change to verified status come from Cloudflare’s announcement, read on September 11. The scope covering existing free-plan customers comes from the company’s press release, which states it more fully. The choices governing how far each block reaches come from Cloudflare’s developer changelog. Google’s objection is reported by TechCrunch and predates this change; what Google-Extended controls comes from Google’s own crawler documentation. The traffic figures and the company’s description of its own footprint are Cloudflare’s, drawn from Cloudflare Radar and its Investor Day 2026 presentation. The closed beta status of Pay Per Crawl comes from its documentation.