The internet just got a lot harder for AI to quietly scrape, and most sites didn't even have to ask

Cloudflare just quietly rewrote the rules on how AI companies are allowed to scrape the internet, and most site owners did not have to lift a finger.
From 15 September, Cloudflare blocks "mixed use" crawlers by default on any page that carries ads. These are bots that blend search indexing with AI training and agent retrieval under one user agent, meaning a single crawler doing double duty for a search engine and a training pipeline. Cloudflare now buckets that traffic into whichever category is most restrictive, effectively shutting the door unless the site owner deliberately opens it back up.
The change automatically covers new Cloudflare customers, any new site set up by existing customers, and everyone still on the free plan. Existing paid customers keep whatever settings they already had, so this is very much a default shift rather than a blanket ban.
The tricky part is separating traffic that genuinely helps a site, like search indexing, from traffic that just feeds a model with no benefit flowing back. Googlebot, Applebot and Bingbot all crawl for both search and training under the same identity, and Cloudflare cannot cleanly tell those two jobs apart at the request level. The new default effectively forces AI companies to choose, or negotiate directly with publishers for the training access they want.
It is a quiet infrastructure change rather than a headline grabbing showdown, but the flow on effect is significant. A meaningful chunk of the open web that AI models have trained on for free is about to get a lot harder to reach without a licence or a deal in place.




Comments