TechnologySeptember 18, 2026

When News Became Training Data: What the Unsealed Filings Say

Key Vocabulary

scrape/skreɪp/
To copy large amounts of data from websites for collection or analysis.
"Engineers scrape web pages to gather training text."
dataset/ˈdeɪ.tə.sɛt/
A structured collection of data used for training or analysis.
"The dataset included millions of news documents."
paywall/ˈpeɪ.wɔːl/
A technical or business barrier that restricts access to online content unless paid for.
"Researchers discussed methods to bypass a paywall."
substitutive/səbˈstɪt.jʊ.tɪv/
Acting as a replacement for something else, especially original work.
"Plaintiffs argued the outputs were largely substitutive."
infringement/ɪnˈfrɪndʒ.mənt/
The violation of a law or a right, especially copyright.
"The suit accuses companies of copyright infringement."

Listening

When News Became Training Data: What the Unsealed Filings Say

New material unsealed in the consolidated copyright litigation has laid out how OpenAI and Microsoft collected and shared news content for model training. The filings state that OpenAI’s mid-training datasets contained 91,692 copies of works from the New York Times, the Daily News and the Center for Investigative Reporting, and that a Common Crawl-derived dataset included more than 2 million documents from nytimes.com. They also describe internal transfers of data under names such as Project Taxi and Project Mango, and the Project Mango collection is said to include at least 160,903 unique works from the publishers.

Internal presentations and emails have warned that widespread scraping could create a damaging doom loop and that generative systems might substitute for the reporting that produced the training material. Microsoft director of applied science Brent Hecht described the scale of unlicensed copying as an astonishing theft and warned of a risk to the employment of people who generated the content. OpenAI leaders, including the head of ChatGPT, have noted that the models are strong at news tasks and that the products can be largely substitutive, which plaintiffs say undermines a fair use defense.

The unsealed brief, filed with the New York court in early September 2026, also offers examples of technical steps that researchers discussed, such as plans to bypass paywalls and to remove copyright notices from training files. Company testimony has been taken; Microsoft chief executive Satya Nadella has said that paywalled material should be licensed, while the filings show OpenAI staff discussed ways to limit outputs tied to plaintiffs. The news organizations seek rulings in a motion for summary judgment and have argued that the scale and the internal admissions should decide liability before trial.

283 words

Quiz

1. Which Microsoft executive has said paywalled material should be licensed?
2. How many unique works did Project Mango contain as cited in filings?
3. What legal filing did the news organizations seek that was unsealed in September 2026?

Reading Practice

Read the article from the Listening section aloud. Your AI teacher will give you pronunciation feedback.

Discussion

1

Do you follow any news outlets closely? Have you stopped visiting some because of summaries?

2

Have you seen AI tools give short news answers instead of links? How did that make you feel?

3

What would you do if your favorite writer lost work because of automation?

4

Have you ever paid for news online? Why or why not?

5

Would you prefer detailed articles or short AI summaries when you need fast information?

此內容僅供英語學習使用,不保證事實的準確性。