Destroying Books to Build a Mind

0 13


Petras, the bookstore owner in Toronto, said that she was generally comfortable selling to undercover buyers, but that she would never sell them a book with interesting marginalia or an important provenance. Joyce Kosofsky, one of the owners of Brattle Book Shop, in Boston, said that the books she sold all had multiple copies. “We probably had them at the cheapest price,” she guessed. In general, she argued, books are no different from any other saleable good: “Just like when you go into a clothing store, and you buy a pair of jeans—they’re your jeans. You can wear them. You can decorate them. You can give them away. No one follows you around saying, ‘What are you going to do with your jeans?’ ”

Given that Anthropic is purchasing books that are “likely not very rare,” Martin, the Princeton professor, said, the company’s use of book guillotines shouldn’t be considered inflammatory. But, if the books aren’t rare, then what are they? The booksellers shared the names of more than six hundred titles that they believed they had sold to A.I. companies, and I sent the list to Melanie Walsh, an Assistant Professor in the Information School at the University of Washington, to process digitally. The texts were obscure in their subject matter and lack of popularity: these were books, sometimes with low-print runs—usually a thousand copies or fewer—to meet realistic market demands. Of the top ten publishers, eight were academic presses—roughly a third of the sample over all. Most of the books were published between the nineteen-seventies and the twenty-tens, and the genres spanned history, biography, fiction, poetry, literary criticism, law, and the social sciences. “Based on this sample, it appears that A.I. companies may be interested in training models on peer-reviewed academic research across a wide range of subjects,” Walsh concluded. This dovetails with what Mycal Tucker, a research scientist at Anthropic, discovered while organizing training data for an A.I. model: “Nonfiction works tend to be more valuable than fiction,” he said in his written testimony, adding that nonfiction-book data helped the model perform well in disciplines as diverse as philosophy and astronomy.

Walsh said that when it comes to nonfiction works that are specialized, as is the case with “Utilization of Municipal Wastewater Sludge,” a 1972 booklet that Petras recently sold—“there’s an argument to be made that these obscure academic books may make a bigger impact as part of a Claude model than they would otherwise.” After all, they were already headed toward obsolescence.

Anthropic’s attempts to get its hands on “all the books in the world” has attracted legal challenges. In 2025, the company agreed to pay $1.5 billion to settle a class-action lawsuit brought by a group of authors who accused the company of violating their copyrights, namely by using their books for A.I. training without their permission. William Alsup, the judge presiding over the case, reprimanded Anthropic for some of its actions, such as downloading over seven million pirated copies of books and keeping the files “as a permanent, general-purpose resource,” even if they weren’t being used to train Claude. (“Anthropic seems to believe that because some of the works it copied were sometimes used in training L.L.M.s, Anthropic was entitled to take for free all the works in the world and keep them forever with no further accounting,” Alsup wrote in his decision.)

But the larger practice of using books to train A.I. models was fair use, according to Alsup. “Authors cannot rightly exclude anyone from using their works for training or learning,” he wrote. “Everyone reads texts, too, then writes new texts. They may need to pay for getting their hands on a text in the first instance. But to make anyone pay specifically for the use of a book each time they read it, each time they recall it from memory, each time they later draw upon it when writing new things in new ways would be unthinkable.”



Source link

Leave A Reply

Your email address will not be published.