MICROSOFT

The Ctrl+C audit is complete

Microsoft says Copilot rarely repeats large chunks of news articles or books, as it fights copyright lawsuits from publishers and authors.

As part of the case, Microsoft shared 8.2 million Copilot chats with an expert working for news publishers.

These chats were picked because they were more likely than usual to include material from the publishers involved.

Microsoft says 59,545 chats contained at least 16 words that matched news content used to help Copilot answer questions.

An expert for the Center for Investigative Reporting found 51 cases with a larger amount of matching text.

In a separate lawsuit involving book authors, Microsoft says only 24 Copilot responses contained at least 30 matching words.

Out of 212 books checked, only 10 had any matches.

The New York Times disagrees with Microsoft’s view of the results.

Its lawyers argue that Microsoft and OpenAI used copyrighted journalism to build AI products that now compete with the publishers whose work was used.

Here’s what you should know:

  • Microsoft says Copilot rarely copies large chunks of publishers’ work.

  • Publishers and authors argue the bigger issue is that their work was used to build competing AI products.

  • The case could help decide how copyright law applies to AI training in the US.

Copilot gets Turnitin’d

Microsoft says the figures support its argument that using copyrighted material to train AI should count as fair use.

It argues that AI models use the material for a different purpose, and that the small number of copied passages does not change that.

The lawsuits against Microsoft and OpenAI are now being handled by one judge.

Microsoft is asking the judge to end the case before it goes to a full trial. If that request is rejected, the legal fight will continue.

Microsoft went digging for plagiarism and came back with pocket lint.- MG