TechnologySeptember 18, 2026

When News Became Training Data: What the Unsealed Filings Say

Key Vocabulary

scrape/skreɪp/
To copy large amounts of data from websites.
"They scrape web articles to build a dataset."
training/ˈtreɪnɪŋ/
Teaching a computer model using data.
"The model needs training on many articles."
paywall/ˈpeɪ.wɔːl/
A system that limits access to online content unless you pay.
"The article was behind a paywall."

Listening

When News Became Training Data: What the Unsealed Filings Say

Unsealed court papers show Microsoft and OpenAI employees worried about how A.I. systems use news stories. The New York Times and other publishers filed the documents in a copyright case.

A Microsoft scientist, Brent Hecht, wrote that the scraping of news was the largest theft of labor in human history. The filings say OpenAI and Microsoft used millions of news articles to train models.

Microsoft data shows its Copilot answer engine cut click-through rates to The New York Times by as much as 93%. The filings also say staff discussed ways to share training data in projects called Project Mango and Project Taxi.

103 words

Quiz

1. Who wrote that the scraping of news was the largest theft of labor in human history?
2. Which product did Microsoft data say cut click-through rates to The New York Times by as much as 93%?
3. Which news organization filed the documents in the copyright case?

Reading Practice

Read the article from the Listening section aloud. Your AI teacher will give you pronunciation feedback.

Discussion

1

Do you read news online? How do you feel when a site is behind a paywall?

2

Have you ever used an AI chat tool to get news? What did you notice?

3

What do you think about computers learning from many news articles?

此内容仅供英语学习使用,不保证事实的准确性。