{"id":614778,"date":"2026-09-05T00:10:10","date_gmt":"2026-09-05T00:10:10","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/614778\/"},"modified":"2026-09-05T00:10:10","modified_gmt":"2026-09-05T00:10:10","slug":"attribution-decay-complicates-the-picture-of-ai-generated-images-scientists-find-the-art-newspaper","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/614778\/","title":{"rendered":"\u2018Attribution decay\u2019 complicates the picture of AI-generated images, scientists find &#8211; The Art Newspaper"},"content":{"rendered":"<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">Art generated with artificial intelligence (AI) relies on models trained on billions of images. That training has been the subject of much debate and several <a class=\"transition-colors duration-default shadow-externalLink hover:text-red-1\" href=\"https:\/\/jipel.law.nyu.edu\/andersen-v-stability-ai-the-landmark-case-unpacking-the-copyright-risks-of-ai-image-generators\/\" target=\"_blank\" rel=\"nofollow noopener\">ongoing<\/a> or impending <a class=\"transition-all duration-default shadow-internalLink hover:text-red-1\" href=\"https:\/\/www.theartnewspaper.com\/2024\/05\/10\/deviantart-midjourney-stable-diffusion-artificial-intelligence-image-generators\" rel=\"nofollow noopener\" target=\"_blank\">lawsuits<\/a>. Many consider the training to be a form of involuntary extraction and that the images derived from it are a violation of countless artists\u2019 copyrights, but recent scientific findings about how AI processes the information it is trained on complicate this view.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">Scientists at the Massachusetts Institute of Technology\u2019s Computer Science and Artificial Intelligence Laboratory published new research in <a class=\"transition-colors duration-default shadow-externalLink hover:text-red-1\" href=\"https:\/\/www.nature.com\/articles\/s41467-026-75667-5\" target=\"_blank\" rel=\"nofollow noopener\">Nature Communications<\/a> on 18 August that explores a phenomenon they have termed \u201cattribution decay\u201d. The researchers, Zheng Dai and David K. Gifford, found that in large datasets, removing certain data did not change the output. The argument that could be made in light of this goes something like: if a copyrighted image was part of the training data, and removing it did not change the generated image, then the generated image could not be accused of copyright infringement.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">The researchers \u201cdeveloped a method for taking away one piece of the training sets and then regenerating the image as though that piece of training data didn\u2019t exist\u201d, Dai says. \u201cAnd if you find that [the output] doesn\u2019t change much, then you can\u2019t attribute it to that piece of data, because it didn\u2019t have any influence on the final output.\u201d He adds that in the context of image generation, &#8220;when you train on large data sets, there is no piece of data that you can take out that significantly alters the image, and that [leads] to this idea of unattributability\u201d.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">So for example, if A Bigger Splash (1967) and even David Hockney\u2019s entire oeuvre were removed from a dataset, could the AI tool still churn out a Bigger Splash-like image with the right set of prompts? Dai says that such scenarios would be speculative, but if the training data contained traces of A Bigger Splash, the AI tool could still generate something resembling the famous painting at Tate Britain. In other words, even if all direct references to A Bigger Splash are removed, the dataset might still include enough indirect references, derivative images like parodies, glimpses of it or echoes of it to generate a Bigger Splash-like image. In this hypothetical scenario, Hockney&#8217;s original painting would not be part of the dataset the generation relied on, so one could argue in court that there was no copyright infringement.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">\u201cIf you view [the] training off of people\u2019s data as [infringement],\u201d Dai says, \u201cit is sort of ironic, the more infringement you do, the less infringing the output is.\u201d<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">Framing their findings as scientific evidence that theft is OK, so long as it is done at a massive scale, is unlikely to win anyone over to the pro-AI side of the argument. Although AI tools can allow for direct infringement when prompted to replicate specific source material, the public sentiment against such models does not account for the sheer magnitude of the datasets involved. When these algorithms generate images based on broader prompts, they are referencing datasets comprising millions if not billions of images. To use a crude metaphor, if every online image of an artwork created by an artist were a grain of sand, to accuse an AI artist of violating your copyright each time they generate an image is a bit like walking up to a stranger on a beach and accusing them of building a sandcastle with your sand.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">A loophole in this clumsy metaphor is that the systems at work are capable of choosing precisely which grains of sand to use if prompted to do so. More importantly, there is a big difference between holding an AI user accountable for the outputs their prompts generate, and holding the AI company accountable for how they trained their model.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">Asked if his and Gifford\u2019s findings effectively serve as a get-out-of-jail-for-free card for tech companies being accused of copyright infringement over their AI models, Dai has a caveat. \u201cIt certainly looks like it from one angle, but I think it is trickier than that,\u201d he says. \u201cFor example, if every single artist joined a class action lawsuit and said, \u2018Oh no, you can\u2019t use [our works]\u2019&#8230; If you take out every single piece of artwork [from a dataset], then obviously the model wouldn&#8217;t exist.\u201d But, he added that in a legal battle, companies could remove a copyrighted image to show its absence does not change the output, demonstrating unattributability as an argument against copyright infringement.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">It seems unlikely that a global artists\u2019 movement will lobby to collectively opt out of all AI training, but even if we did, AI models would still be able to draw from works that are in the public domain or derivative works, and a plethora of additional types of material we may not have even considered. Dai and Gifford\u2019s findings do not exonerate the training process itself; even if an AI company were to argue for unattributability in a particular output, it could still be held liable for how its algorithms were trained.<\/p>\n<p class=\"pt-dp-p font-text-light font-light text-lg leading-normal tracking-wide mb-base last:mb-0\" itemprop=\"text\">Attribution decay complicates an already contentious legal landscape in which the public perception is often that all generated output violates copyright, and none of it is copyrightable. To be sure, AI-generated works of art can still be copyrighted under certain conditions. The United States Copyright Office insists on a <a class=\"transition-all duration-default shadow-internalLink hover:text-red-1\" href=\"https:\/\/www.theartnewspaper.com\/2023\/05\/04\/us-copyright-office-artificial-intelligence-art-regulation\" rel=\"nofollow noopener\" target=\"_blank\">case-by-case approach<\/a> that looks for the existence of human creative expression.<\/p>\n","protected":false},"excerpt":{"rendered":"Art generated with artificial intelligence (AI) relies on models trained on billions of images. That training has been&hellip;\n","protected":false},"author":2,"featured_media":614779,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[260754,218,61,3546,60,80],"class_list":["post-614778","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology","tag-art-technology","tag-artificial-intelligence","tag-ie","tag-intellectual-property","tag-ireland","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/614778","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=614778"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/614778\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/614779"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=614778"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=614778"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=614778"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}