When News Became Training Data: What the Unsealed Filings Say
Key Vocabulary
Listening
When News Became Training Data: What the Unsealed Filings Say
New material unsealed in the consolidated copyright litigation has laid out how OpenAI and Microsoft collected and shared news content for model training. The filings state that OpenAI’s mid-training datasets contained 91,692 copies of works from the New York Times, the Daily News and the Center for Investigative Reporting, and that a Common Crawl-derived dataset included more than 2 million documents from nytimes.com. They also describe internal transfers of data under names such as Project Taxi and Project Mango, and the Project Mango collection is said to include at least 160,903 unique works from the publishers.
Internal presentations and emails have warned that widespread scraping could create a damaging doom loop and that generative systems might substitute for the reporting that produced the training material. Microsoft director of applied science Brent Hecht described the scale of unlicensed copying as an astonishing theft and warned of a risk to the employment of people who generated the content. OpenAI leaders, including the head of ChatGPT, have noted that the models are strong at news tasks and that the products can be largely substitutive, which plaintiffs say undermines a fair use defense.
The unsealed brief, filed with the New York court in early September 2026, also offers examples of technical steps that researchers discussed, such as plans to bypass paywalls and to remove copyright notices from training files. Company testimony has been taken; Microsoft chief executive Satya Nadella has said that paywalled material should be licensed, while the filings show OpenAI staff discussed ways to limit outputs tied to plaintiffs. The news organizations seek rulings in a motion for summary judgment and have argued that the scale and the internal admissions should decide liability before trial.
Quiz
Reading Practice
Read the article from the Listening section aloud. Your AI teacher will give you pronunciation feedback.
Discussion
Do you follow any news outlets closely? Have you stopped visiting some because of summaries?
Have you seen AI tools give short news answers instead of links? How did that make you feel?
What would you do if your favorite writer lost work because of automation?
Have you ever paid for news online? Why or why not?
Would you prefer detailed articles or short AI summaries when you need fast information?