AI Firms Acquire Vast Quantities of Pre-2022 Books for Training Data
Artificial intelligence companies are reportedly purchasing large volumes of books printed before 2022. These older works are sought after because they are considered "AI-free" data, meaning they were not generated or influenced by AI systems. Developers are acquiring these books on a massive scale, with the intention of scanning them. During the scanning process, the books are often destroyed. This practice raises questions concerning copyright and the preservation of literary works.
The acquisition of pre-2022 books highlights a critical bottleneck in the development of advanced AI models: the need for vast, diverse, and ethically sourced training datasets. As AI capabilities grow, the demand for unique data that avoids potential copyright entanglements or AI-generated biases intensifies. This trend suggests a potential market for digitized older literature, but also raises concerns about the destruction of physical artifacts and the long-term implications for cultural heritage. Future AI development may need to balance data acquisition with preservation efforts and explore more sustainable methods for data generation and curation.
AI-generated to prompt reflection — not editorial opinion, not advice, not a statement of fact. How this works.