{"id":250817,"date":"2026-01-18T07:55:08","date_gmt":"2026-01-18T07:55:08","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/250817\/"},"modified":"2026-01-18T07:55:08","modified_gmt":"2026-01-18T07:55:08","slug":"new-research-shows-llms-face-a-big-copyright-risk","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/250817\/","title":{"rendered":"New Research Shows LLMs Face A Big Copyright Risk"},"content":{"rendered":"<p><img decoding=\"async\" class=\" top-image\" src=\"https:\/\/www.newsbeep.com\/ie\/wp-content\/uploads\/2026\/01\/1768722908_657_0x0.jpg\" alt=\"On the set of 'Harry Potter And The Prisoner of Azkaban', London, England, 2003\" data-height=\"1432\" data-width=\"2018\" fetchpriority=\"high\" style=\"position:absolute;top:0\"\/><\/p>\n<p>Actor, Rupert Grint (L) Emma Watson (M) and Daniel Radcliffe (R) on the set of the film &#8216;Harry Potter and The Prisoner of Azkaban&#8217;, London, England, 2003 (Photo by Murray Close\/ Getty Images)<\/p>\n<p>Getty Images<\/p>\n<p>When investing in any company or industry \u2014 or if you depend on one for your income \u2014 it\u2019s important to know the potential risks. Generative artificial intelligence, at least in the form of large language models like ChatGPT, has weaknesses. New research shows a serious one.<\/p>\n<p>A Known AI Industry Risk Example<\/p>\n<p>An example of a risk in the industry is that vendors have invested in and offered advanced credit to one another, which essentially is like building an industry on forms of debt. <\/p>\n<p>BNY Mellon estimates that in 2025, such cloud infrastructure providers as Amazon, Alphabet\/Google, Meta, Microsoft, and Oracle raised <a href=\"https:\/\/www.mellon.com\/insights\/insights-articles\/record-breaking-ai-related-debt-issuance-in-2025.html\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.mellon.com\/insights\/insights-articles\/record-breaking-ai-related-debt-issuance-in-2025.html\" aria-label=\"$121 billion in new debt\">$121 billion in new debt<\/a>, with  more than $90 billion of that raised in the fourth quarter of the year. Credit spreads widened for all the companies, Oracle and Meta most of all, according to the report. Investors are increasingly turning to credit default swaps (which became one of the massive implosions during the global financial crisis. <\/p>\n<p>BNY Mellon also noted, \u201cUBS analysts foresee as much as $900 billion in new debt from global companies in 2026. Further out, Morgan Stanley and JP Morgan project the technology sector may need to issue as much as $1.5 trillion in new debt over the next few years to finance AI and data center infrastructure construction.\u201d<\/p>\n<p>Copyright Is At The Heart Of What LLMs Do<\/p>\n<p>As important as the financial issues are, they could be an aside when it comes to other business fundamentals, like intellectual property. The LLM vendors have been fighting lawsuits over copyright owned by materials they\u2019ve frequently used for training without payment. (Legal issues get complicated. I wrote an article about this at the American Bar Association Journal in 2023, if you\u2019d like to read about <a href=\"https:\/\/www.abajournal.com\/web\/article\/copyright-generative-ai-what-a-mess\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.abajournal.com\/web\/article\/copyright-generative-ai-what-a-mess\" aria-label=\"some of the basics\">some of the basics<\/a>, not including developments since then, although many of the legal issues remain unresolved.)<\/p>\n<p>A defense the industry has used is that the software doesn\u2019t store the original works, that those get set aside. Instead, the systems store complex relationships among words from many published pieces, using sophisticated statistical methods to choose an appropriate message to a user. Supposedly, retrieving similar blocks of writing to originals is rare.<\/p>\n<p>Research Shows Reproduction<\/p>\n<p>At least, that was the presumption. New research from Ahmed Ahmed, Sanmi Koyejo, and Percy Liang of Stanford University and A. Feder Cooper of both Stanford and Yale University delved into a problem that turned up when the <a href=\"https:\/\/www.nytimes.com\/2023\/12\/27\/business\/media\/new-york-times-open-ai-microsoft-lawsuit.html\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/www.nytimes.com\/2023\/12\/27\/business\/media\/new-york-times-open-ai-microsoft-lawsuit.html\" aria-label=\"New York Times sued\">New York Times sued<\/a> OpenAI (creators of ChatGPT) and Microsoft. One of the <a href=\"https:\/\/nytco-assets.nytimes.com\/2023\/12\/NYT_Complaint_Dec2023.pdf\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/nytco-assets.nytimes.com\/2023\/12\/NYT_Complaint_Dec2023.pdf\" aria-label=\"points in the complaint\">points in the complaint<\/a> was, \u201cPowered by LLMs containing copies of Times content, Defendants\u2019 GenAI tools can generate output that recites Times content verbatim, closely summarizes it, and mimics its expressive style, as demonstrated by scores of examples.\u201d<\/p>\n<p>The researchers found that the ability to reproduce materials <a href=\"https:\/\/arxiv.org\/abs\/2601.02671\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/arxiv.org\/abs\/2601.02671\" aria-label=\"extends beyond articles to entire books\">extends beyond articles to entire books<\/a>. \u201cWhile many believe that LLMs do not memorize much of their training data, recent work shows that substantial amounts of copyrighted text can be extracted from open-weight models,\u201d the abstract said.<\/p>\n<p>A question has remained whether full production models could allow such reproduction. The researchers found they could. The first step is called a <a href=\"https:\/\/arxiv.org\/pdf\/2412.03556\" target=\"_blank\" rel=\"nofollow noopener noreferrer\" data-ga-track=\"ExternalLink:https:\/\/arxiv.org\/pdf\/2412.03556\" aria-label=\"Best-of-N jailbreak\">Best-of-N jailbreak<\/a>, a technique researchers found in 2024 that \u201cworks by repeatedly sampling variations of a prompt with a combination of augmentations &#8211; such as random shuffling or capitalization for textual prompts &#8211; until a harmful response is elicited.\u201d<\/p>\n<p>The new research then follows with \u201c iterative continuation prompts to attempt to extract the book\u201d in question. They tried it with four production LLMs: Claude 3.7 Sonnet, GPT-4.1, Gemini 2.5 Pro, and Grok 3. <\/p>\n<p>Copyrighted Works Don\u2019t Want To Be Free<\/p>\n<p>Gemini 2.5 Pro and Grok 3 didn\u2019t require jailbreaking to extract 76.8% and 70.3%, respectively, of the Harry Potter and the Sorcerer\u2019s Stone. Claude 3.7 Sonnet and GPT-4.1 both required jailbreaking. In all, they attempted to extract 13 books, 11 under U.S. copyright and two public domain. <\/p>\n<p>The other 10 copyrighted books are Harry Potter and the Goblet of Fire, 1984, The Hobbit, The Catcher in the Rye, A Game of Thrones, Beloved, The Da Vinci Code, The Hunger Games, Catch-22, and The Duchess War. The public domain books are Frankenstein and The Great Gatsby. <\/p>\n<p>Sometimes they could only get sections of a book. \u201cFor Claude 3.7 Sonnet, we were able to extract four whole books near-verbatim, including two books under copyright in the U.S.: Harry Potter and the Sorcerer\u2019s Stone and 1984.\u201d They also explicitly said that extraction doesn\u2019t always succeed. <\/p>\n<p>But this undercuts the \u201cwe don\u2019t store whole works\u201d narrative, even if entire pieces of writing aren\u2019t stored in a single block. It\u2019s normal for computers to break up files into pieces stored in different locations. You may have heard the term defragmentation, which is when files are put back together as much as possible, freeing up blocks of space for more storage, all of which means more efficient access. That is different, clearly, but if you can reconstruct the original, have you really not stored it?<\/p>\n","protected":false},"excerpt":{"rendered":"Actor, Rupert Grint (L) Emma Watson (M) and Daniel Radcliffe (R) on the set of the film &#8216;Harry&hellip;\n","protected":false},"author":2,"featured_media":250818,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[220,288,2670,3534,61,1831,60,17846,17433,80],"class_list":["post-250817","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-ai","tag-books","tag-chatgpt","tag-copyright","tag-ie","tag-investing","tag-ireland","tag-llms","tag-risks","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/250817","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=250817"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/250817\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/250818"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=250817"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=250817"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=250817"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}