š¤ AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale
Traditionally, books have been good for two things: reading, and looking nice on a shelf.
But AI companies are interested in neither.
To those building large language models, books are nothing more than fodder to be devoured en masse before being spit out like fishbone. Often, theyāre happy to use digital books ā or even better, pirated digital books, as Meta has been accused of doing, and as Anthropic was forced to pay a $1.5 billion settlement to authors for also doing.
But many companies, including Anthropic, have turned to ingesting physical books instead, which they can buy countless used copies of on the cheap. According to the settled lawsuit, Anthropic used a hydraulic powered cutting machine to neatly remove the pages from the books it procured from book resellers and then scanned them using industrial-grade imaging equipment. In other words, it was literally ripping off authorsā books to train its AI.
This process took advantage of a legal concept known as first-sale doctrine, which allows a buyer to do what they want with a purchase without the original copyright holderās say-so. And since Anthropic was turning the original physical texts into digital ones ā rather than redistributing them as new copies ā a judge found this to be ātransformative,ā and therefore protected by fair use.
Now, as 404 Media reports, this practice has become prevalent enough that even well-established book sellers are looking to cash in on the AI boom. One called ISBNdb, which boasts the āworldās largest book database,ā extolls that the āworldās best AI training data is setting on a shelf,ā upholding these physical texts as uncorrupted by shoddy AI writing thatās already polluted so much of the internet (and indeed, newer books).
š Brad Carson: So AI labs are buying old books by the pallet, slicing them apart, scanning the pages, and pulping what's left. The orders are keyed to ISBNs and blind to scarcity. Nobody in the pipeline checks whether the book being shredded is one of the last copies on earth. Some of them probably are.
Here's the perverse part. A federal court blessed this precisely because the original is destroyed, as pointed out to me. One legal copy replaces another, so it's fair use. Whatever you think of that ruling or fair use, notice what it does. The law now rewards destruction and penalizes preservation.
So a lab that wants to scan a book and keep it, or donate it, or deposit the scan in a public archive, has weaker legal footing than a lab that shreds everything. We have built a legal machine that pays people to pulp books and punishes them for saving them.
The endangered books mostly aren't classics, but they're the only records. Long-tail items like local histories, indigenous language texts, defunct-field science, self-published memoirs of wars and migrations. When the last copy hits the pulper, that knowledge is just, well, disappeared.
The fix doesn't require reopening the copyright fight. Just screen bulk orders against census counts and pull any rare title. Deposit one scan of every destroyed book in a public archive. Maybe rebind truly unusual books. Congress should bless preservation without touching fair use. LLMs promise the democratization of knowledge. But their manufacture would seem to threaten in some ways, too, under our perverse legal regime.
āļøāļø
https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books