{"id":580735,"date":"2026-08-02T20:11:12","date_gmt":"2026-08-02T20:11:12","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/580735\/"},"modified":"2026-08-02T20:11:12","modified_gmt":"2026-08-02T20:11:12","slug":"why-ai-companies-are-buying-millions-of-books-to-destroy-them","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/580735\/","title":{"rendered":"Why AI companies are buying millions of books to destroy them"},"content":{"rendered":"<p>AI companies are buying and then destroying millions of rare books to prevent <a href=\"https:\/\/www.jpost.com\/tags\/ai-industry-analysis\" target=\"_blank\" rel=\"nofollow noopener\">AI slop<\/a> content that consumers are vocally against, 404 Media reported last week.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">Silicon Valley tech giants are paying companies and contractors to buy up rare books, which are then scanned in a high-speed machine that cuts their spines out and then shreds the originals.<\/p>\n<p>The tech giants are reportedly buying up the rare <a href=\"https:\/\/www.jpost.com\/tags\/books\" target=\"_blank\" rel=\"nofollow noopener\">books<\/a> to train new AI models and prevent AI \u201cslop.\u201d In one article on its site, ISBNdb, a company that claims to have the \u201cthe world\u2019s largest book database,\u201d argued that books published before 2022 were best for AI training data because they would not have any AI-generated text.<\/p>\n<p>In a report uncovered by the Washington Post in January, one Anthropic co-founder suggested that feeding AI models books could teach them \u201chow to write well\u201d instead of producing \u201clow-quality internet speak.\u201d<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cThe world&#8217;s best AI training data is sitting on a shelf,\u201d ISBNdb wrote in a since-deleted blog post. \u201cPrint books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [&#8230;] \u00a0\u201cPhysical books published before this date [pre-2022] are structurally clean of modern poisoning tools.\u201d<\/p>\n<p><img alt=\"FILE PHOTO: Anthropic logo is seen in this illustration taken May 20, 2024.\" loading=\"lazy\" width=\"822\" height=\"829\" decoding=\"async\" data-nimg=\"1\" style=\"color:transparent\" src=\"https:\/\/images.jpost.com\/image\/upload\/f_auto,fl_lossy\/c_fill,g_faces:center,h_537,w_822\/709424\"\/>FILE PHOTO: Anthropic logo is seen in this illustration taken May 20, 2024. (credit: REUTERS\/DADO RUVIC\/ILLUSTRATION\/FILE PHOTO)<\/p>\n<p>404 Media reported that the AI tech companies were interested in buying up books to prevent \u201cmodel collapse,\u201d where AI models trained on lower-quality AI-generated results progressively lose quality.\u00a0 Executives believed that vast troves of books were essential to give the models new data and prevent their model collapse.<\/p>\n<p>Why are AI companies destroying old books?<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">Notably, special edition book sellers say that they usually sell one or two books to a single customer.<\/p>\n<p>\u201cIn the rare book trade, it\u2019s very seldom that people want to buy more than one book,\u201d antique book seller Pieter de Vries told The Telegraph in an interview published last week. \u201cSo if somebody comes and says, \u2018I want a couple of hundred of your books,\u2019 it\u2019s very strange.\u201d<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">But ISBNdb and companies like it are now helping AI tech giants purchase orders ranging from 1,000 to one million books.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">The AI companies\u2019 attempt to hoover up books, art, news articles, and other forms of media has gotten attention in several instances since 2024.<\/p>\n<p>Most notably, the Washington Post reported in January on Anthropic\u2019s attempts to buy millions of books, slice their spines, and scan their pages to feed more data into the company\u2019s chatbot, Claude.<\/p>\n<p><img alt=\"Old books.\" loading=\"lazy\" width=\"822\" height=\"829\" decoding=\"async\" data-nimg=\"1\" style=\"color:transparent\" src=\"https:\/\/images.jpost.com\/image\/upload\/f_auto,fl_lossy\/c_fill,g_faces:center,h_537,w_822\/732382\"\/>Old books. (credit: Jan Mellstr\u00f6m\/ Stock image)Anthropic&#8217;s legal battle over AI and book copyrights<\/p>\n<p>That instance later led to a multi-million dollar class action lawsuit in which several authors sued <a href=\"https:\/\/www.jpost.com\/business-and-innovation\/all-news\/article-899253\" target=\"_blank\" rel=\"nofollow noopener\">Anthropic<\/a>, arguing that the company, which is \u200cbacked by \u2060Amazon and Alphabet (Google&#8217;s parent company), used pirated versions of their books without permission to teach Claude to respond to human prompts.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">Anthropic settled with the authors at $1.5 billion.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">Notably, though, the fact that Anthropic destroyed the books made the company\u2019s case stronger.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">Judge William Alsup ruled last June that Anthropic made fair \u2060use of the authors&#8217; work to train Claude, but found that the company violated their rights by saving more than seven million pirated books to a &#8220;central library&#8221; that would not necessarily be used for AI training.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cHere, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,\u201d Alsup wrote.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cThe print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company,\u201d he said.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">In short, because the AI company pulped the books, the digital version that existed afterward \u201creplaced\u201d the physical one. The judge ruled that the companies merely transformed the books, which meant that it was not a violation of US copyright law.<\/p>\n<p>The tech giants clearly wanted this secret since the beginning, because the optics of destroying books are unpalatable for many. Documents uncovered by the Washington Post found that even internally, companies wanted to distance themselves from the effort.<\/p>\n<p>\u201cProject Panama is our effort to destructively scan all the books in the world,\u201d one internal document unsealed in legal filings last said, as reported by the Washington Post. \u201cWe don\u2019t want it to be known that we are working on this.\u201d<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">As ISBNdb wrote in a since-deleted blog post on its website: \u201cThe optics problem is real. \u2018AI company destroys two million books\u2019 is not a headline that generates sympathy.\u201d<\/p>\n<p>Tech giants, former officials decry AI companies for destroying books<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">But the headlines and lawsuits became public, and the public is rather unsympathetic.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cSo AI labs are buying old books by the pallet, slicing them apart, scanning the pages, and pulping what&#8217;s left,\u201d former US House rep. Brad Carson wrote in a post on X\/Twitter.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">&#8220;Here&#8217;s the perverse part. A federal court blessed this precisely because the original is destroyed. One legal copy replaces another, so it&#8217;s fair use. Whatever you think of that ruling or fair use, notice what it does. The law now rewards destruction and penalizes preservation.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cA lab that wants to scan a book and keep it, or donate it, or deposit the scan in a public archive, has weaker legal footing than a lab that shreds everything. We have built a legal machine that pays people to pulp books and punishes them for saving them.\u201d<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">Some tech giants have voiced their distaste for the destruction of the books.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cI\u2019ve asked the SpaceX AI team to preserve any rare books in a library and scan them the hard way rather than just cutting off the spine and scanning,\u201d Elon Musk said in an X\/Twitter post.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">&#8220;There&#8217;s something particularly misanthropic about the mechanized destruction of such intimate human objects,&#8221; said SEO of Factory AI Matan Grinberg. &#8220;History seldom looks kindly on those who destroy books, whatever the reasons.&#8221;<\/p>\n<p>One critic of the AI industry\u2019s approach to copyrighted work and the founder of Fairly Trained, a creator-rights group, Ed Newton-Rex, told the Telegraph that the secrecy with which companies like ISBNdb and Anthropic acquire books is damning.<\/p>\n<p>\u201cClearly both the provider of these books and the AI companies know that this is a terrible look and they don\u2019t want the specifics to get out,\u201d he told the Telegraph.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cIf you are just going and spending $1 on a used book, with all of the money going to a book wholesaler, should that give you the right to train a commercial generative AI model on that book, which will then be able to compete with the author who wrote it?\u201d Newton-Rex said. \u201cA lot of people, myself included, think it shouldn\u2019t.\u201d<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cThere is surely no more fitting image in the generative AI age for the exploitation that underlies this technology than almost trillion-dollar companies buying books for a few cents or a dollar each, scanning them, training on them, then destroying them, essentially subsuming culture,\u201d he added.<\/p>\n<p>In response to the flurry of reporting around the destroyed books, ISBNdb disputed the reports that it helped purchase large volumes of books for <a href=\"https:\/\/www.jpost.com\/breaking-news\/article-820026\" target=\"_blank\" rel=\"nofollow noopener\">AI training<\/a>.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cWe&#8217;ve seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised,\u201d the company wrote in a statement.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cThe facts: ISBNdb has never purchased, scanned, or sold a book &#8211; for AI training or anything else. We don&#8217;t train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We&#8217;ve taken the page down.<\/p>\n<p class=\"article-paragraph-section article-body-paragraph\">\u201cOur job is helping people find books. For more than two decades, ISBNdb has been the card catalog of the book world &#8211; the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn&#8217;t changed.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"AI companies are buying and then destroying millions of rare books to prevent AI slop content that consumers&hellip;\n","protected":false},"author":2,"featured_media":580736,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[220,218,219,288,61,60,80],"class_list":["post-580735","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-books","tag-ie","tag-ireland","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/580735","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=580735"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/580735\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/580736"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=580735"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=580735"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=580735"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}