AI August 2, 2026 bearish ⇧ 56 pts across 2 threads

AI Crawlers Take and Never Give Back

Two threads converged on the same ugly math. Only 8.9% of sites block AI crawlers, but 94.8% of sites are never cited in AI answers. A commenter nailed the implication immediately: it is architecturally impossible for an LLM to associate a crawled link with a generated response, so 'citations' are a cosmetic feature, not a real attribution system. The data was framed as an SEO problem, but it is really a property rights problem.

The rare books thread made the same point from a different angle. AI firms are physically destroying rare book editions after scanning them, which some commenters defended as a net positive for access. Others called it a convenient workaround for copyright complaints: you can't violate copyright on a book that no longer exists. Both threads share a core anxiety: content is being consumed at scale, and the people who made it get nothing in return, not traffic, not credit, not compensation.

The pattern here is that the web's implicit deal, you publish, people visit, you benefit, is broken. Whether your content is a rare 18th-century manuscript or a blog post, the AI pipeline treats it the same way: raw material with no return address.


So what?

If you run a content business or any site where organic traffic matters, blocking AI crawlers is now a serious strategic question, not a technical afterthought. The data says most sites don't block crawlers and most don't get cited anyway, so the current default (allow crawling, hope for citations) is delivering almost nothing. Founders building on top of web content should assume the citation economy will not materialize and plan accordingly.

Read these