Cloudflare is set to make a major change to huge portions of the internet after standing between websites and automated traffic in a bid to protect publishers and hosters. Soon, the company will divide AI-related traffic into one of three categories – search, agent or training – in a bid to give customers more control over just how much automated traffic they’re prepared to let in.
Beginning September 2026, new Cloudflare domains will block agent and training crawlers by default on pages that display ads, arguing that publishers need human visitors to see those ads to generate a revenue.
But huge challenges stand between now and such a system being rolled out effectively, because many major crawlers like Googlebot are defined as multipurpose and don’t exactly fit into one of those three categories. In the worst case scenario, this could prevent any automated traffic from reaching a site, even if some categories are permitted.
Cloudflare vs. Google: A battle in which publishers could be the actual losers
Worryingly, Googlebot specifically plays a key role in discovering and updating pages in Google Search, and preventing its access could have major implications on content indexing.
Pages that reject Googlebot on the basis that publishers wanted to block at least one of those three categories could gradually lose search visibility, leading to major commercial implications.
Despite the clear disadvantages of blocking automated traffic, many assert that blocking crawlers helps to protect their intellectual property, reduces web hosting infrastructure demands caused by unpredictable traffic spikes, and gives owners more control over how and where their information is used.
Varn Search Marketing Search & Innovation Director Andy Mollison describes the situation as a “high-stakes staring contest” between Cloudflare and Google, but more importantly, one in which publishers are actually the ones who pull the short straw.
I spoke with Mollison to understand how publishers and businesses should move forward to strike the right balance between protecting their content and remaining visible online.
Andy, when I came across this email, the first thing that came to my mind is that maybe Google wants it to happen to teach a (very harsh) lesson to those who block AI bots. A bit of a "Squid Game" moment but for websites?
Cloudflare and Google are currently having a high-stakes staring contest. On paper this looks like a mismatch. And in reality, it doesn’t look like it will work out for Cloudflare.
There is a simple reason Google isn’t going to blink first; the only losers from this will be Cloudflare’s customers. Mainly, publishers.
People will still be using Google, blissfully unaware a significant portion of the web is no longer included in search results. Cloudflare customers that rely on ad revenue, on the other hand will see their traffic drop off a cliff.
The goal is potentially admirable; allowing publishers to protect content from being stolen by AI, but by setting the default to ‘block all AI bots’ publishers that aren’t following every move Cloudflare makes closely will be the ones who pay the price.
What is the worst case scenario for publishers, B2B businesses, consumer brands and ultimately consumers?
The worst case scenario is every website with ads currently using Cloudflare will be completely blocked from Google search.
This obviously poses a problem for publishers in terms of the organic traffic loss. The knock on effect is then decreased ad revenue. Similarly for B2B businesses and consumer brands, this lack of search visibility will harm their bottom lines.
Consumers may not initially notice, but eventually they might question why results from publishers stop appearing.
Cloudflare, in its Content Independence Day, explained the logic behind the decision while still giving content publishers sovereignty over their decisions. Why is Cloudflare taking this stance?
It is the answer to why most businesses do most things, but the simple answer is money.
Cloudflare’s CEO declared last month that for the first time in the internet’s history, machines now generate more web traffic than people. AI crawlers are a big reason why. They take up a lot of crawl budget on servers, and therefore Cloudflare needs more resources to handle this increased traffic.
Cost is probably the primary reason, but a secondary aspect is autonomy. Cloudflare wants to give content publishers the ability to decide what level of indexation their content receives, depending on the intention of the AI robot.
For example, you could block your content from being put into an AI's training model, but still keep it available for when someone performs a 'search' in an AI platform.
This gives publishers better choice on where their content is used in AI platforms, on a page by page basis.
You mentioned mixed messages from Google? What are they and why are these so difficult to decipher?
A little bit technical here, but bear with me. Cloudflare has been quite persistent with trying to recommend an llms.txt website file, which could in theory be used to dictate which parts of your website are open to AI crawlers. And more importantly, which parts are not.
Google's stance has been to state that an llms.txt file is not something they recommend in their Developer Fundamentals guidance. However, they have also created pages in their Agentic Browsing documentation about how to set up an llms.txt file.
So the confusion comes in Google stating they don't use llms.txt files, yet still creating documentation for using it.
There feels like there is a misalignment between Google's developers, which in itself then becomes difficult for website owners to know which stance to take. Either Google can crawl without use of the file, or they need it. It can’t be both.
Why would brands, publishers and anyone that owns a website want to block AI traffic?
A significant portion of searches now ends without a click. For pu...