Uno dei motivi per cui non mi è mai piaciuto Cloudflare è che pensano di sapere cosa è meglio per te e i default sono spesso dannosi (es. caching).
Ora hanno deciso di loro iniziativa di bloccare buona parte dei crawler AI tirando però dentro incidentalmente, se non ho capito male, anche i bot dei motori di ricerca che permettono la normale indicizzazione:
On September 15, 2026, we’ll be setting new defaults for each of these three classifications. For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.
[...] Another change that will apply on September 15 is that multi-purpose crawlers (specifically those that combine Search with Training) will be allowed/blocked according to all of their behaviors, in line with our call for transparency for website owners. Since the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training (either through the new options to manage AI traffic, or through the legacy Block AI bots service).
Spero di aver capito male.
EDIT: la parafrasi che fanno i media è "Cloudflare sets deadline to block AI crawlers that bundle search with AI training". La scadenza sarebbe quindi un modo per costringere le "aziende AI" a differenziare i crawler tra ricerca, training e uso agentico.