Cloudflare pushes AI crawlers to separate search from training
Cloudflare will block crawlers that mix search, training and agent activity by default on ad-supported pages. It is also developing a model that pays publishers when their content creates value in AI answers.
Summary
Cloudflare is changing its AI crawler rules to draw a clearer line between search, model training and agent use. From September 15, 2026, crawlers that combine several of these functions will be blocked by default on ad-supported pages when they do not let site owners choose between each purpose.
The change will apply to new customers, new sites created by existing customers and current free-plan customers who have not changed their settings beforehand. Site owners will still be able to adjust these options in their Cloudflare dashboard.
In practice
Search will remain allowed by default, while training and agent access will be blocked on pages displaying ads. Cloudflare argues that a crawler using the same identity for search and training forces publishers to choose between discoverability and allowing their content to be used for other purposes.
At the same time, the company is evolving Pay Per Crawl into Pay Per Use. The aim is to compensate a publisher when its content is used in an answer or accessed as premium material, rather than merely when a page is fetched. Initial tests involve Ceramic.ai and You.com.
What we still don't know
Cloudflare controls a significant part of the Web's infrastructure, but it does not set the rules for every site or AI company. The policy's effectiveness will depend on accurate crawler identification, publisher adoption and whether AI platforms separate their different uses transparently.
Pay Per Use is still at an early stage. Cloudflare has not announced universal prices, payment levels or revenue guarantees for publishers, and outcomes may differ substantially between major publications and smaller sites.
Why it matters
- It makes it harder to treat search visibility as automatic permission to use content for training or agents.
- Publishers gain more specific control over who can collect and reuse their pages.
- It tests an alternative to licensing deals negotiated mainly by large media groups.
- AI companies and search engines may face greater pressure to state the purpose of each crawler clearly.