Court documents unsealed Thursday in the New York Times' copyright suit against OpenAI and Microsoft reveal that Microsoft staffers questioned whether AI scraping represented “the largest theft of labor in human history.” A 2023 Microsoft memo warned that millions of people would view large models “hoovering up” their work as “an astonishing theft of unprecedented proportions.” Microsoft said the document was written by Brent Hecht, a director of applied science, and did not reflect company policy.
Microsoft CEO Satya Nadella testified that “anything that is paywalled should be licensed by anyone who wants to use it,” and said he would have exercised Microsoft's right to force OpenAI to retrain models if he had known paywalled content was used. Internal OpenAI exchanges included a staffer telling president Greg Brockman about building a “hack” to bypass the NYT paywall, to which Brockman replied: “ah nice.” Nick Turley, who ran the ChatGPT team, wrote that AI posed an “existential threat” to publishers, while former policy director Jack Clark warned as early as 2020 that OpenAI was creating systems that substitute for cultural labor.
Both OpenAI and Microsoft argue training on publicly available material is fair use. Steven Lieberman, representing the New York Daily News and seven other papers, said: “The world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior.” Judge Sidney Stein is weighing summary judgment motions; the case has already forced OpenAI to preserve 20 million ChatGPT conversation logs.
Separately, OpenAI unveiled Astra for Law on September 17, a GPT-6 Astra configuration built around a legal search index of more than 230 million URLs of U.S. case law, statutes, regulations, court rules, and administrative decisions. The launch came exactly one week after OpenAI introduced ChatGPT for Financial Services with data from LSEG, PitchBook, and Daloopa. Astra for Law scored a 54% pass rate on Vals AI's Legal Research Bench, compared with 38.7% for GPT-6 Astra with standard web search, a 40% relative improvement. It also found 24% more reference cases and retrieved up to 54% more relevant passages from correct opinions.
The legal product launched with 26 partner-built plugins and nine community plugins, including Thomson Reuters, Intapp, iManage, Relativity, Clio, and DeepJudge. Harvey Technologies and Legora are API customers, while early law-firm collaborators include Sullivan & Cromwell, Ropes & Gray, Cooley, Latham & Watkins, and Wachtell Lipton. Thomson Reuters CTO Joel Hron said legal professionals need “trusted intelligence, relevant enterprise and matter context, purpose built legal capabilities, and the governance required for high stakes work.”
The seven-day sequence between the finance and law launches underscores how quickly OpenAI is moving into regulated verticals, building its own domain-specific data layers and workflow tools rather than simply licensing data. For data vendors, the speed of that expansion is a direct competitive challenge.