{"id":7704,"date":"2025-09-10T16:44:14","date_gmt":"2025-09-10T16:44:14","guid":{"rendered":"https:\/\/www.newsbeep.com\/il\/7704\/"},"modified":"2025-09-10T16:44:14","modified_gmt":"2025-09-10T16:44:14","slug":"rss-co-creator-launches-new-protocol-for-ai-data-licensing","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/il\/7704\/","title":{"rendered":"RSS co-creator launches new protocol for AI data licensing"},"content":{"rendered":"<p id=\"speakable-summary\" class=\"wp-block-paragraph\">In the wake of Anthropic\u2019s $1.5 billion copyright settlement, the AI industry is coming to terms with its training data problem. There are <a rel=\"nofollow noopener\" href=\"https:\/\/www.theatlantic.com\/technology\/archive\/2025\/07\/anthropic-meta-ai-rulings\/683526\/\" target=\"_blank\">as many as 40 other pending cases<\/a> that seek damages for unlicensed data \u2014 including one that takes Midjourney to court for <a rel=\"nofollow noopener\" href=\"https:\/\/apnews.com\/article\/warner-bros-midjourney-ai-copyright-lawsuit-dc-studios-b87d80d7b4a4dfdcf0ee149d30830551\" target=\"_blank\">creating images of Superman<\/a>. <\/p>\n<p class=\"wp-block-paragraph\">Without some kind of licensing system, AI companies could face an avalanche of copyright lawsuits that <a rel=\"nofollow noopener\" href=\"https:\/\/www.lawfaremedia.org\/article\/anthropic-s-settlement-shows-the-u.s.-can-t-afford-ai-copyright-lawsuits\" target=\"_blank\">some worry<\/a> will set the industry back permanently.<\/p>\n<p class=\"wp-block-paragraph\">Now, a group of technologists and web publishers has launched a system that would enable data licensing at massive scale \u2014 provided AI companies take them up on it. Called Real Simple Licensing (RSL), the system is already being backed by major web publishers like Reddit, Quora and Yahoo. The question now is if that momentum will be enough to bring major AI labs to the bargaining table.<\/p>\n<p class=\"wp-block-paragraph\">According to RSL co-founder Eckart Walther, who also co-created the RSS standard, the goal was to create a training-data licensing system that could scale across the internet. \u201cWe need to have machine-readable licensing agreements for the internet,\u201d Walther told TechCrunch. \u201cThat\u2019s really what RSL solves.\u201d<\/p>\n<p class=\"wp-block-paragraph\">For years, groups like the Dataset Providers Alliance have been pushing for clearer collection practices, but RSL is the first attempt at a technical and legal infrastructure that could make it work in practice. On the technical side, <a rel=\"nofollow noopener\" href=\"https:\/\/rslstandard.org\/\" target=\"_blank\">the RSL Protocol<\/a> lays out specific licensing terms a publisher can set for their content, whether that means AI companies need a custom license or to adopt Creative Commons provisions. Participating websites will include the terms as part of their \u201crobots.txt\u201d file in a prearranged format, making it straightforward to identify which data falls under which terms.<\/p>\n<p class=\"wp-block-paragraph\">On the legal side, the RSL team has established a collective licensing organization, <a rel=\"nofollow noopener\" href=\"https:\/\/rslcollective.org\" target=\"_blank\">the RSL Collective<\/a>, that can negotiate terms and collect royalties, similar to ASCAP for musicians or MPLC for films. As in music and film, the goal is to give licensors a single point of contact for paying royalties, and provide rightsholders a way to set terms with dozens of potential licensors at once.<\/p>\n<p class=\"wp-block-paragraph\">A host of web publishers have already joined the collective, including Yahoo, Reddit, Medium, O\u2019Reilly Media, Ziff Davis (owner of Mashable and Cnet), Internet Brands (owner of WebMD), People Inc. and The Daily Beast. Others, like Fastly, Quora and Adweek, are supporting the standard without joining the collective.<\/p>\n<p>Techcrunch event<\/p>\n<p>\n\t\t\t\t\t\t\t\t\tSan Francisco<br \/>\n\t\t\t\t\t\t\t\t\t\t\t\t\t|<br \/>\n\t\t\t\t\t\t\t\t\t\t\t\t\tOctober 27-29, 2025\n\t\t\t\t\t\t\t<\/p>\n<p class=\"wp-block-paragraph\">Notably, the RSL Collective includes some publishers that already have licensing deals \u2014 most notably Reddit, which receives <a rel=\"nofollow noopener\" href=\"https:\/\/www.reuters.com\/technology\/reddit-ai-content-licensing-deal-with-google-sources-say-2024-02-22\/\" target=\"_blank\">an estimated $60 million a year<\/a> from Google for use of its training data. There\u2019s nothing stopping companies from cutting their own deals within the RSL system, just as Taylor Swift can set special terms for licensing while still collecting royalties through ASCAP. But for publishers too small to draw their own deals, RSL\u2019s collective terms are likely to be the only option.<\/p>\n<p class=\"wp-block-paragraph\">But while it\u2019s easy enough to determine when a song has been played, AI models pose unique challenges when it comes to figuring out when royalties are due for a specific piece of training data. The issue is simplest for a product like Google\u2019s AI Search Abstracts, which draw data from the web in real time and maintain strict attribution for each fact. <\/p>\n<p class=\"wp-block-paragraph\">But if training isn\u2019t logged when it occurs, it can be nearly impossible to confirm that a given document was ingested into a LLM. It\u2019s particularly challenging if publishers ask to be paid per-inference rather than receiving a blanket fee, an option offered by one of the stock RSL licenses.<\/p>\n<p class=\"wp-block-paragraph\">Still, RSL\u2019s creators believe AI companies will be able to manage the difficulty. \u201cSome of the licensing agreements they\u2019ve already done have required them to be able to report on it, so it\u2019s possible,\u201d says Doug Leeds, a co-founder of RSL and former CEO of IAC Publishing. \u201cIt doesn\u2019t have to be perfect. It just has to be good enough to get people paid.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The bigger question is whether AI companies will embrace the system. As the success of companies like ScaleAI and Mercor shows, frontier labs have no problem paying for data, but the web has traditionally been seen as a source for cheap, low-quality data. With datasets like the Common Crawl already available, it may be a challenge to extract royalties from something labs are used to getting for free. And as <a rel=\"nofollow noopener\" href=\"https:\/\/blog.cloudflare.com\/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives\/\" target=\"_blank\">the recent dustup<\/a> between CloudFlare and Perplexity shows, it\u2019s not straightforward to tell the difference between web-scraping and machine-enhanced browsing.<\/p>\n<p class=\"wp-block-paragraph\">When I put the question to Leeds, he pointed to recent comments from AI leaders calling for a system like RSL \u2014 most notably <a rel=\"nofollow noopener\" href=\"https:\/\/www.mrjonathanjones.com\/2024\/12\/15\/shaping-the-ai-content-frontier-deals-data-and-value-exchange\/#:~:text=become%20standard%20practice.-,Creators%20Designing%20Content%20for%20AI%20Systems,specifically%20to%20improve%20AI%20systems.&amp;text=Succinctly%20capturing%20the%20idea%20of,sets%20for%20emerging%20generative%20systems.\" target=\"_blank\">from Sundar Pichai at last year\u2019s Dealbook Summit<\/a>. Whether the calls for a licensing system are earnest or not, the RSL team plans to hold them to it. \u201cThey have said outwardly to everyone, something like this needs to exist,\u201d Leeds told me. \u201cWe need a protocol. We need a system.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Now, they may get one.<\/p>\n","protected":false},"excerpt":{"rendered":"In the wake of Anthropic\u2019s $1.5 billion copyright settlement, the AI industry is coming to terms with its&hellip;\n","protected":false},"author":2,"featured_media":7705,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[345,8945,343,344,8946,85,46,8947,8948,125],"class_list":["post-7704","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-ai-copyright","tag-artificial-intelligence","tag-artificialintelligence","tag-eckart-walther","tag-il","tag-israel","tag-real-simple-licensing","tag-rsl","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/7704","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/comments?post=7704"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/posts\/7704\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media\/7705"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/media?parent=7704"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/categories?post=7704"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/il\/wp-json\/wp\/v2\/tags?post=7704"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}