Cloudflare Restricts Access for AI Crawlers

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
The New Rules for AI Crawler Agents
AI crawler agents, those automated robots that explore the web to retrieve real-time information, will now need to obtain permission to access certain parts of the web. This measure, announced by Cloudflare on July 1, will take effect on September 15. Until now, these bots could freely explore the web, but from now on, they will be blocked by default on certain pages. This change has sparked considerable discussion, particularly regarding its impact on Google, a major player in this field.
Cloudflare has introduced a new classification for bots, divided into three distinct categories. The first, Search, pertains to bots that index pages to respond to subsequent queries. The second, Agent, includes automated systems that act in real-time for a user, such as ChatGPT retrieval bots. Finally, the Training category encompasses crawlers that collect content to train AI models. These new rules have been in effect since July 1 for all Cloudflare customers, including those using the free service.
Upcoming Changes and Implications
Starting September 15, Cloudflare's default settings will change. The Training and Agent categories will be blocked on pages containing advertisements, while Search will remain permitted. These new restrictions will apply to new domains integrated into Cloudflare, new sites created by existing customers, as well as all free-tier customers. Those wishing to avoid these restrictions will need to adjust their security settings before this date.
The rationale behind this decision is based on the idea that the presence of advertisements on a page indicates it is intended for a human audience. A search crawler that redirects a user to a page is considered beneficial, while a bot that reads the page and relays the information to a third party is not.
Challenges for AI Crawler Agents
AI agents were developed with the assumption that the open web would remain accessible. For instance, a search agent can retrieve a competitor's pricing, a monitoring tool can check a supplier's listings, and a customer service agent can extract product specifications. Until now, these actions did not require a license, but that situation is changing.
Cloudflare, which manages a significant portion of global web traffic, is implementing these blocks at the network level, making them harder to circumvent than a simple directive in a robots.txt file. The pages supported by advertising are precisely those that agents seek to explore, as they contain crucial information such as news, reviews, prices, and product descriptions. For a business, the failure of an agent does not result in a lawsuit but rather in silence or an incomplete response based on the information still accessible.
A notable issue concerns Google. The Googlebot, used for both search and training, could be blocked under the new rules. Cloudflare CEO Matthew Prince expressed hope that these changes would encourage companies to separate the search and training functions of their bots, thereby highlighting the pressure to comply with the new rules.
How to Obtain Permission for Agents
Agent managers must first identify which of their Cloudflare accounts will be classified as Agent. This classification is based on the behavior of the bots rather than a formal registration. A real-time browsing search agent will therefore be affected, even if its operator does not consider it a crawler.
Content publishers also have decisions to make. They need to check their service level, as free-tier customers will automatically be subject to the new default settings starting September 15. They must also assess whether blocking the Training category is worth the potential cost, as this could also affect the Googlebot and, consequently, their visibility in search results.
A key aspect to watch is monetization. The pay-per-use model, with companies like Ceramic.ai and You.com compensating publishers when their content is used by AI agents, could become the norm. Cloudflare noted that more than half of AI crawler traffic is dedicated to retrieving unchanged pages, which represents a waste of resources on both sides.
This change marks the beginning of a new era in online content management, where the proposed solution is a pricing model rather than a simple access restriction. However, the classification of bots into Search, Agent, and Training relies on the companies' own declarations, which could prompt some to reclassify their activities to avoid restrictions. The free and unlimited access to the web, which has lasted for thirty years, is now being called into question, and those who do not adapt quickly may face significant obstacles.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.