Amazon, once an online bookseller, is destroying rare books to train AI models
Rare books are incredibly valuable for training LLMs, since these models have already trained on whatever's available online.
Amazon's decision to destroy rare books for training AI models highlights the increasing demand for unique and high-quality data to train large language models (LLMs). As LLMs have already been trained on vast amounts of online data, they require more specialized and rare content to continue improving. This move by Amazon, which has evolved from an online bookseller to a tech giant, underscores the value of rare and out-of-print books in advancing AI research.
The destruction of rare books for AI training purposes raises questions about the preservation of cultural heritage and the ethics of using unique materials for commercial purposes. However, it also highlights the growing need for data curation and preservation in the AI era. As AI models continue to drive innovation in various industries, the importance of high-quality training data will only continue to grow. This trend is likely to lead to increased investment in data curation, preservation, and creation.
What's next to watch is how other companies and organizations respond to the growing demand for rare and unique data. Will we see a rise in data curation services, or new business models emerge for preserving and providing access to rare materials? Additionally, as AI models become increasingly important, will there be a shift towards more transparent and explainable AI training practices, and how will regulators and policymakers respond to these developments?
Originally reported by techcrunch.com. IndexNews adds analysis for ai & agent economy readers.