We’ve long joked about delaying bedtime in order to read the entire internet, but the prospect of actually doing so doesn’t sound as appealing as it used to. An ever-growing proportion of text online, we may sense, hasn’t been written by humans, but generated by artificial intelligence. One solution is to return to reading on paper, and beyond that, spending our time only with books published before the advent — and thus uncorrupted by the output — of large language models. Ironically, that’s what large language models themselves have had to do, and in order to ensure them a steady diet of manmade text, their owners are resorting to controversially destructive methods.
Suspicions first arose when book dealers noticed a spike in their income due to large orders from mysterious buyers for titles they assumed they’d never unload. Books on home oxygen treatment in Italy, marriage litigation in medieval England, Texas civil procedure in the twenty-tens, Swedish comedy in the sixties: these entities bought them all, and for whatever price the seller happened to be asking.
404 Media reporter Emanuel Maiberg tracked the path of one such book, which eventually wound up at a Las Vegas scanning facility run by Amazon. Though undergirded by advanced technology, the work done there is simple: digitize as many books as possible, then destroy them.
“In order to scan them really fast, they cut the spines off the books and fed the pages to the scanner,” Maiberg explained to NPR. Many are technically rare, albeit esoteric obscurities like vanity publications and old instruction manuals rather than first-edition classics. Humans may not value them, but LLMs do: their consumption forestalls “a problem called model collapse, where if an AI trains on AI training data, it gets worse.” Amazon and Anthropic have both been found to use the book-sacrificing method to harvest text, but every company with models to train needs a solution. All this makes for “an unusually on-the-nose example of digital technology devouring literary culture,” as the Wall Street Journal’s Melissa Korn puts it. But when you really have read the entire internet, perhaps extreme measures are inevitable.
Related Content:
How the Internet Archive Digitizes 3,500 Books a Day–the Hard Way, One Page at a Time
How 99% of Ancient Literature Was Lost
Based in Seoul, Colin Marshall writes and broadcasts on cities, language, and culture. He’s the author of the newsletter Books on Cities as well as the books 한국 요약 금지 (No Summarizing Korea) and Korean Newtro. Follow him on the social network formerly known as Twitter at @colinmarshall.
Leave a Reply