{"id":851445,"date":"2026-09-17T22:54:10","date_gmt":"2026-09-17T22:54:10","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/851445\/"},"modified":"2026-09-17T22:54:10","modified_gmt":"2026-09-17T22:54:10","slug":"microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/851445\/","title":{"rendered":"Microsoft exec called AI scraping \u2018the largest theft of labor in human history,&#8217; new unredacted filings reveal"},"content":{"rendered":"<p id=\"speakable-summary\" class=\"wp-block-paragraph\">New unredacted information in <a href=\"https:\/\/techcrunch.com\/2023\/12\/27\/the-new-york-times-wants-openai-and-microsoft-to-pay-for-training-data\/\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">the copyright lawsuit The New York Times<\/a> brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications.<\/p>\n<p class=\"wp-block-paragraph\">Per the lawsuit, a top Microsoft executive privately described the companies\u2019 AI training practices as \u201ctheft,\u201d and OpenAI\u2019s own leadership said its AI models posed an \u201cexistential threat\u201d to the publishers and journalists whose work trained them.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The unsealed material also details how the companies allegedly obtained and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">It\u2019s worth noting that much of the new information comes from The Times\u2019 own brief, not the underlying exhibits, which remain sealed. The quotes below are presented without their original context. <\/p>\n<p class=\"wp-block-paragraph\">The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times initially alleged the firms violated copyright law by training generative AI models on its content.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies\u2019 arguments that training constitutes \u201cfair use.\u201d This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the <a href=\"https:\/\/techcrunch.com\/2026\/09\/02\/u-s-government-sides-with-openai-on-issue-of-training-llms-on-copyrighted-material\/\" rel=\"nofollow noopener\" target=\"_blank\">Trump administration contributed a brief<\/a> in defense of OpenAI\u2019s unlicensed use of copyrighted material to train its LLMs.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Several of the new admissions, however, run counter to OpenAI\u2019s fair use defense, particularly the rule\u2019s requirement that use doesn\u2019t substitute or harm the market for the original work.<\/p>\n<p class=\"wp-block-paragraph\">For example, Microsoft\u2019s own data shows its Copilot \u201canswer engine\u201d caused click-through rates for The New York Times\u2019 domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft\u2019s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a \u201cdoom loop\u201d that would \u201churt the performance of our models and the entire web at the same time.\u201d<\/p>\n<p class=\"wp-block-paragraph\">\u201cIt is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its \u2018content supply chain,\u2019\u201d reads the Microsoft document, as quoted in the filing.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Microsoft CEO Satya Nadella also testified in a deposition earlier this year that \u201canything that is paywalled should be licensed by anyone who wants to use it\u2026for grounding or training,\u201d and made clear that, if he \u201chad been made aware that OpenAI had scraped and trained on information that was behind a paywall,\u201d he would have \u201cinvoked [Microsoft\u2019s right to] require OpenAI to retrain its models.\u201d\u00a0 <\/p>\n<p class=\"wp-block-paragraph\">Other admissions cut against different pillars of the fair-use test: OpenAI\u2019s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an \u201cexistential threat\u201d from products like the chatbot, which are \u201clargely substitutive\u201d and \u201cwill get more and more substitutive as they get better.\u201d<\/p>\n<p class=\"wp-block-paragraph\">OpenAI President Greg Brockman described the models as \u201cexcellent at news.\u201d Nadella agreed under oath earlier this year that conversing with chatbots \u201chas substituted \u2026 giving you the information right there on the website on the AI platform versus needing to go to the underlying source.\u201d<\/p>\n<p class=\"wp-block-paragraph\">That kind of language speaks to how the technology could directly compete with, rather than transform, the original work.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">A Microsoft document states that there is a \u201creal risk\u201d that generative AI could \u201csignificantly disrupt the employment of the very people who generated the data on which the foundation model was trained.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI\u2019s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from <a href=\"http:\/\/nytimes.com\" target=\"_blank\" rel=\"noreferrer noopener nofollow\">nytimes.com<\/a> alone.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In a January 2023 internal memo, Hecht called it \u201can astonishing theft of unprecedented proportions\u201d and \u201cthe largest theft of labor in human history.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The filing lays out in new detail how OpenAI and Microsoft went about acquiring the plaintiffs\u2019 content, including scraping it from the Bing Index.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cOpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI\u2019s models within its own commercial products,\u201d the filing reads. \u201cMicrosoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The companies allegedly assembled the Project Mango data into a training dataset that contains copies of at least 160,903 unique works from the news publishers.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In order to get the most out of their scraping, OpenAI employees allegedly came up with a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder told Brockman about a \u201chack to get around nytimes paywall,\u201d Brockman replied: \u201cah nice.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">OpenAI employees also allegedly built training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly pulled millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers \u201cwouldn\u2019t want model outputting\u201d \u201ccopyright notices\u201d to users.<\/p>\n<p class=\"wp-block-paragraph\">\u201cThe evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,\u201d Steven Lieberman, counsel for the New York Daily News, said in a statement shared with TechCrunch. <\/p>\n<p class=\"wp-block-paragraph\">OpenAI and Microsoft did not return requests for comment.<\/p>\n<p>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" rel=\"nofollow noopener\" target=\"_blank\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/p>\n","protected":false},"excerpt":{"rendered":"New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years&hellip;\n","protected":false},"author":2,"featured_media":851446,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[27],"tags":[28,23517,1281,19721,1283],"class_list":["post-851445","post","type-post","status-publish","format-standard","has-post-thumbnail","category-business","tag-business","tag-copyright","tag-microsoft","tag-new-york-times","tag-openai"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/851445","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=851445"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/851445\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/851446"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=851445"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=851445"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=851445"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}