AI training data demand is distorting book markets
A thread about a 'bananas' order for 5,000 obscure book titles is treating it as an obvious AI training data acquisition play. The books being ordered include outdated editions ('Pass Your Driving Test, 2018 Edition'), which fits the pattern of scraping for volume rather than quality. One commenter calls it 'the Paperclip Maximiser in the form of the Training Material Maximiser,' which is vivid and probably accurate.
The pattern: AI training data demand is now large enough to visibly distort secondary markets for physical goods. This is the book equivalent of what GPU demand did to gaming hardware. The commenter suggestion that these companies should digitize the books via Google Books-style scanning rather than buying physical copies suggests the inefficiency is real and probably intentional (harder to detect, harder to litigate against).
The practical implication for anyone in publishing, archiving, or content licensing is that your back catalog has new and unexpected value.
So what?
If you have a content library, an archive, or access to out-of-print material, AI training data demand is a real monetization channel right now. If you are building AI products, the supply chain for training data is getting more visible and more contested, which means licensing costs and legal risk are both going up.