Cloudflare's AI crawler controls are mostly theater so far
Cloudflare announced new AI traffic options for customers, letting site owners block or allow specific AI crawlers. The HN reaction was blunt: it's still the honors system. Bots that don't respect robots.txt won't respect Cloudflare's new settings either. The more interesting buried detail in the thread is that one of the 'block training' policies will block Googlebot starting September 15th, because Google uses the same crawler infrastructure for both search indexing and AI training. That is not a small thing.
The pattern: the web's relationship with AI scrapers is entering a genuinely adversarial phase. Site owners want control, but the tools are lagging. Cloudflare is offering the interface before the enforcement. The Google crawler overlap issue is particularly sharp because it means choosing to block AI training could mean choosing to drop out of Google's search index, which is an impossible tradeoff for most sites.
There's also no update on Cloudflare's pay-per-crawl program, which was supposed to let AI companies compensate publishers for crawling. The silence on that suggests it's not moving fast.
So what?
If you run a content site and are considering blocking AI training crawlers, the Google crawler overlap issue means you may face a real SEO cost. Don't configure these settings without understanding which crawlers share infrastructure. The pay-per-crawl model remains vaporware for now, so don't build revenue assumptions around it.