Skip to content

AI companies buy and destroy old books for ‘clean’ training data

“I don’t like that uncommon books are being pulped.”

Summarised by Centrist

AI companies are buying physical books in bulk, cutting off their bindings and scanning the pages to train large language models.

Anthropic previously used a hydraulic cutting machine and industrial scanners to digitise millions of purchased books. A US judge found the process transformative and protected by fair use because the company was not distributing replacement copies.

Now ISBNdb, which describes itself as the world’s largest book database, is helping anonymous buyers order between 1,000 and one million books at a time.

“The world’s best AI training data is sitting on a shelf,” the company says.

Books published before the rise of generative AI are particularly valuable because their contents are uncontaminated by AI-generated material increasingly found online.

“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination,” ISBNdb said.

The company also promises confidentiality, acknowledging that “AI company destroys two million books” would not generate sympathetic headlines.

Booksellers in several countries have reported unexplained bulk orders, including random selections of rare, foreign-language and out-of-print books. Buyers’ identities and intended uses are generally concealed, meaning sellers can only suspect AI laboratories are responsible.

One bookseller said the orders cleared unwanted inventory and provided welcome income, but added: “I don’t like that uncommon books are being pulped.”

The concern is that AI training could destroy some of the few surviving copies of obscure works, preserving their contents inside proprietary models while removing the original books from public circulation.

Read more over at Futurism

Receive our free newsletter here

Latest