Why AI Companies Are Buying—And Then Destroying—Old Books

We’ve long joked about delay­ing bed­time in order to read the entire inter­net, but the prospect of actu­al­ly doing so does­n’t sound as appeal­ing as it used to. An ever-grow­ing pro­por­tion of text online, we may sense, has­n’t been writ­ten by humans, but gen­er­at­ed by arti­fi­cial intel­li­gence. One solu­tion is to return to read­ing on paper, and beyond that, spend­ing our time only with books pub­lished before the advent — and thus uncor­rupt­ed by the out­put — of large lan­guage mod­els. Iron­i­cal­ly, that’s what large lan­guage mod­els them­selves have had to do, and in order to ensure them a steady diet of man­made text, their own­ers are resort­ing to con­tro­ver­sial­ly destruc­tive meth­ods.

Sus­pi­cions first arose when book deal­ers noticed a spike in their income due to large orders from mys­te­ri­ous buy­ers for titles they assumed they’d nev­er unload. Books on home oxy­gen treat­ment in Italy, mar­riage lit­i­ga­tion in medieval Eng­land, Texas civ­il pro­ce­dure in the twen­ty-tens, Swedish com­e­dy in the six­ties: these enti­ties bought them all, and for what­ev­er price the sell­er hap­pened to be ask­ing.

404 Media reporter Emanuel Maiberg tracked the path of one such book, which even­tu­al­ly wound up at a Las Vegas scan­ning facil­i­ty run by Ama­zon. Though under­gird­ed by advanced tech­nol­o­gy, the work done there is sim­ple: dig­i­tize as many books as pos­si­ble, then destroy them.

“In order to scan them real­ly fast, they cut the spines off the books and fed the pages to the scan­ner,” Maiberg explained to NPR. Many are tech­ni­cal­ly rare, albeit eso­teric obscu­ri­ties like van­i­ty pub­li­ca­tions and old instruc­tion man­u­als rather than first-edi­tion clas­sics. Humans may not val­ue them, but LLMs do: their con­sump­tion fore­stalls “a prob­lem called mod­el col­lapse, where if an AI trains on AI train­ing data, it gets worse.” Ama­zon and Anthrop­ic have both been found to use the book-sac­ri­fic­ing method to har­vest text, but every com­pa­ny with mod­els to train needs a solu­tion. All this makes for “an unusu­al­ly on-the-nose exam­ple of dig­i­tal tech­nol­o­gy devour­ing lit­er­ary cul­ture,” as the Wall Street Jour­nal’s Melis­sa Korn puts it. But when you real­ly have read the entire inter­net, per­haps extreme mea­sures are inevitable.

Relat­ed Con­tent:

How the Inter­net Archive Dig­i­tizes 3,500 Books a Day–the Hard Way, One Page at a Time

Libraries & Archivists Are Dig­i­tiz­ing 480,000 Books Pub­lished in 20th Cen­tu­ry That Are Secret­ly in the Pub­lic Domain

How Will AI Change the World?: A Cap­ti­vat­ing Ani­ma­tion Explores the Promise & Per­ils of Arti­fi­cial Intel­li­gence

How to Read Many More Books in a Year: Watch a Short Doc­u­men­tary Fea­tur­ing Some of the World’s Most Beau­ti­ful Book­stores

How 99% of Ancient Lit­er­a­ture Was Lost

Based in Seoul, Col­in Marshall writes and broad­casts on cities, lan­guage, and cul­ture. He’s the author of the newslet­ter Books on Cities as well as the books 한국 요약 금지 (No Sum­ma­riz­ing Korea) and Kore­an Newtro. Fol­low him on the social net­work for­mer­ly known as Twit­ter at @colinmarshall.


by | Permalink | Comments (0) |

Sup­port Open Cul­ture

We’re hop­ing to rely on our loy­al read­ers rather than errat­ic ads. To sup­port Open Cul­ture’s edu­ca­tion­al mis­sion, please con­sid­er mak­ing a dona­tion. We accept Pay­Pal, Ven­mo (@openculture), Patre­on and Cryp­to! Please find all options here. We thank you!


Leave a Reply

Quantcast