Friday, 18 September 2026 Newsarchy UK live index
NewsarchyUKUK
Every UK story. Mapped, sourced, and explained where it matters.
BREAKING
Business

Microsoft executive called AI scraping largest theft of labor in filings

Newly unsealed court documents from an ongoing copyright lawsuit reveal internal Microsoft communications describing AI data scraping as labor theft.

Text:
Microsoft executive called AI scraping largest theft of labor in filings
Microsoft executive called AI scraping largest theft of labor in filings
EXECUTIVE BRIEF Key Takeaways & Signal
  • Core Development: Newly unsealed court documents from an ongoing copyright lawsuit reveal internal Microsoft communications describing AI data scraping as labor theft.
  • Beat Context: Categorized under Business with independent corroboration.
  • Reporting Depth: 3 minute analytical read synthesized from verified newsroom sources.

Newly unsealed court documents from an ongoing copyright lawsuit in Manhattan federal court have revealed stark internal acknowledgments by technology executives, including a senior Microsoft director who privately described artificial intelligence data scraping as the "largest theft of labor in human history."

The disclosures surface amid protracted litigation involving major news organizations, including TechCrunch reporting on the unredacted filings from the case originally launched by New York Daily News and Futurism coverage. The lawsuit contends that tech companies unlawfully harvested millions of copyrighted articles without permission or compensation to build generative artificial intelligence models that now directly compete with the original human labor that supplied their training data.

Media additions

Image via Futurism
Image via Futurism
Image via New York Daily News
Image via New York Daily News
Image via 404 Media
Image via 404 Media

According to the unredacted motion for summary judgment, Microsoft’s Director of Applied Science, Brent Hecht, wrote in an internal memo that large-scale data harvesting would be viewed globally as "an astonishing theft of unprecedented proportions." Hecht warned that the practices threatened the economic stability of content creators, writing that the industry had created a situation where an end-product undermined its own essential suppliers.

Additional documents highlighted by 404 Media show that internal teams at Microsoft tracked how their own artificial intelligence tools cannibalized traffic from foundational web sources. An internal presentation characterized this dynamic as a "doom loop" that degrades both model performance and the wider web simultaneously. Furthermore, Microsoft CEO Satya Nadella conceded in a deposition that paywalled material should be licensed for model training, stating he would have required OpenAI to retrain its models had he known paywalls were bypassed during data collection.

The filings also outline specific technical methods used to amass training corpuses. As reported by GIGAZINE, OpenAI researchers shared methods for circumventing news paywalls undetected, while developers systematically stripped copyright notices from training sets to prevent models from generating attribution. Corresponding dataset figures revealed millions of scraped articles drawn from major publishers across multiple training initiatives.

Publisher / SourceDocumented Impact or Copied Volume
The New York TimesUp to 93% drop in search referral click-through rates on Bing; millions of copies utilized in training datasets.
Daily News and Affiliated PapersMillions of constituent works harvested; survey respondents exhibited higher reliance on AI answers over direct subscriptions.

Legal representation for the media organizations argued that these admissions entirely undermine the tech companies' legal defense of "fair use." Conversely, Microsoft distanced itself from Hecht's remarks in statements provided to the press, asserting that the comments reflected an individual employee's perspective rather than official corporate policy or legal analysis. Representatives for the tech firms maintain that their training methods constitute transformative use under copyright law.

The Manhattan federal court case is expected to proceed toward trial milestones, with judges weighing whether the accumulation and utilization of copyrighted articles without licensing agreements exceeds the legal boundaries of fair use. Additional developments and formal judicial determinations in the copyright litigation are anticipated as the proceedings advance.

READER INTELLIGENCE PULSE

How significant is this development?

Contribute your assessment to the aggregated reader sentiment ledger.

Frequently Asked Questions

Key questions answered in this report

What is the key development in: Microsoft executive called AI scraping largest theft of labor in filings?

Newly unsealed court documents from an ongoing copyright lawsuit reveal internal Microsoft communications describing AI data scraping as labor theft.

Why is this Business development significant for the UK?

This report covers critical events in our Business beat. Independent reporting monitors related UK statements, regulatory shifts, and public responses as further verified details emerge.

How was this reporting corroborated and verified?

Newsarchy UK compiles and cross-references reporting from primary reporting from GIGAZINE and cross-checked wire reports. All coverage adheres to published editorial standards.

When was this report published?

This briefing was published on September 18, 2026 and is permanently cataloged in the Newsarchy UK Business archives.

Related stories