Microsoft executive called AI scraping largest theft of labor in filings
Newly unsealed court documents from an ongoing copyright lawsuit reveal internal Microsoft communications describing AI data scraping as labor theft.
- Core Development: Newly unsealed court documents from an ongoing copyright lawsuit reveal internal Microsoft communications describing AI data scraping as labor theft.
- Beat Context: Categorized under Business with independent corroboration.
- Reporting Depth: 3 minute analytical read synthesized from verified newsroom sources.
Newly unsealed court documents from an ongoing copyright lawsuit in Manhattan federal court have revealed stark internal acknowledgments by technology executives, including a senior Microsoft director who privately described artificial intelligence data scraping as the "largest theft of labor in human history."
The disclosures surface amid protracted litigation involving major news organizations, including TechCrunch reporting on the unredacted filings from the case originally launched by New York Daily News and Futurism coverage. The lawsuit contends that tech companies unlawfully harvested millions of copyrighted articles without permission or compensation to build generative artificial intelligence models that now directly compete with the original human labor that supplied their training data.
Media additions
According to the unredacted motion for summary judgment, Microsoft’s Director of Applied Science, Brent Hecht, wrote in an internal memo that large-scale data harvesting would be viewed globally as "an astonishing theft of unprecedented proportions." Hecht warned that the practices threatened the economic stability of content creators, writing that the industry had created a situation where an end-product undermined its own essential suppliers.
Additional documents highlighted by 404 Media show that internal teams at Microsoft tracked how their own artificial intelligence tools cannibalized traffic from foundational web sources. An internal presentation characterized this dynamic as a "doom loop" that degrades both model performance and the wider web simultaneously. Furthermore, Microsoft CEO Satya Nadella conceded in a deposition that paywalled material should be licensed for model training, stating he would have required OpenAI to retrain its models had he known paywalls were bypassed during data collection.
The filings also outline specific technical methods used to amass training corpuses. As reported by GIGAZINE, OpenAI researchers shared methods for circumventing news paywalls undetected, while developers systematically stripped copyright notices from training sets to prevent models from generating attribution. Corresponding dataset figures revealed millions of scraped articles drawn from major publishers across multiple training initiatives.
| Publisher / Source | Documented Impact or Copied Volume |
|---|---|
| The New York Times | Up to 93% drop in search referral click-through rates on Bing; millions of copies utilized in training datasets. |
| Daily News and Affiliated Papers | Millions of constituent works harvested; survey respondents exhibited higher reliance on AI answers over direct subscriptions. |
Legal representation for the media organizations argued that these admissions entirely undermine the tech companies' legal defense of "fair use." Conversely, Microsoft distanced itself from Hecht's remarks in statements provided to the press, asserting that the comments reflected an individual employee's perspective rather than official corporate policy or legal analysis. Representatives for the tech firms maintain that their training methods constitute transformative use under copyright law.
The Manhattan federal court case is expected to proceed toward trial milestones, with judges weighing whether the accumulation and utilization of copyrighted articles without licensing agreements exceeds the legal boundaries of fair use. Additional developments and formal judicial determinations in the copyright litigation are anticipated as the proceedings advance.
How significant is this development?
Contribute your assessment to the aggregated reader sentiment ledger.
Frequently Asked Questions
Key questions answered in this reportWhat is the key development in: Microsoft executive called AI scraping largest theft of labor in filings?
Newly unsealed court documents from an ongoing copyright lawsuit reveal internal Microsoft communications describing AI data scraping as labor theft.
Why is this Business development significant for the UK?
This report covers critical events in our Business beat. Independent reporting monitors related UK statements, regulatory shifts, and public responses as further verified details emerge.
How was this reporting corroborated and verified?
Newsarchy UK compiles and cross-references reporting from primary reporting from GIGAZINE and cross-checked wire reports. All coverage adheres to published editorial standards.
When was this report published?
This briefing was published on September 18, 2026 and is permanently cataloged in the Newsarchy UK Business archives.