What changes today

Cloudflare’s new default settings for AI crawlers take effect on 15 September. On pages that display advertising, bots that Cloudflare classifies as used for AI training or as AI agents are blocked by default, while search crawlers remain allowed, according to Cloudflare’s announcement in July. Gulf News reported that the change takes effect today.

The new defaults apply to new domains joining Cloudflare and to new sites set up by existing customers, TechCrunch reported when the policy was announced. Existing free customers who had not changed their AI bot settings before the deadline move to the new defaults too, according to Search Engine Journal. Site owners can change the settings in the Cloudflare dashboard.

Three categories and a strictest rule

Cloudflare now sorts bots by purpose. In its definitions, a search crawler “collects or indexes your content, so it can answer questions about it later”. An agent is “automated behavior that is acting, usually in real time, on a person’s behalf”. A training crawler is one “taking your content to train or fine-tune a model”.

The difficulty is crawlers that do more than one of those things. Cloudflare said its defaults “will be enforced by the most restrictive applicable rules”, and named Googlebot, Applebot and Bingbot as multi-purpose crawlers that are blocked for customers who choose to block training. As Search Engine Journal summarised it, “a crawler that performs both Search and Training will be blocked if a site blocks Training”.

A person scrolling through text on a tablet at a table
Cloudflare wants site owners to separate crawlers that help readers find content from those that consume it. Stock photo. iam hogir · pexels · Pexels License

Search Engine Journal noted that Cloudflare’s block works at the network level, so unlike a robots.txt instruction it does not depend on the crawler choosing to comply. It reported no response from Google.

Why Cloudflare is doing it

The policy is meant to push AI companies to run separate crawlers for search, agents and training, and, TechCrunch reported in July, to push them towards paying publishers for content. “Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge,” chief executive Matthew Prince said then.

Cloudflare’s objective is “not simply to block AI”, Gulf News reported, but to let website owners distinguish between automated services that help people discover their content and those that consume it.

A laptop showing a website on a desk
Site owners can change the settings in the Cloudflare dashboard. Stock photo. Thirdman · pexels · Pexels License

The risk for publishers

For an ad-funded site that has chosen to block training, the strictest-rule policy means the main crawlers of Google, Microsoft and Apple can be shut out along with the training bots, which puts search visibility at stake. As long as those companies run crawlers that combine search with training, publishers on Cloudflare face a choice between allowing training and risking search traffic.

What to watch

The pressure now sits with the operators of mixed-use crawlers. The signs to look for are whether Google, Microsoft or Apple split search from training into separate crawlers, and whether Cloudflare publishes figures on how many sites the new defaults cover.