SyloSpace

Why is USA Today suing OpenAI over training data?

News publishers are suing OpenAI over using their journalism to train AI, and the outcome could decide whether that training is legal.

Updated 2 hours ago5 min readVersion 2
CommentsFollow

Covers: Covers the USA Today lawsuit against OpenAI, including the publishers involved, the copyright and related claims, the training data at issue, and the broader context of news publishers suing AI companies. Does not provide legal advice or predict the final outcome of the case.

3 free full reads left this month. Join or upgrade

Business newspaper article
Photo: AbsolutVision

The short answer

Interpretation AI-prepared starting map

USA Today's publisher is among the news organizations suing OpenAI over the use of its journalism as training data for generative AI models. The dispute sits inside a broader wave of copyright litigation between publishers and AI developers, including The New York Times Company's case against OpenAI and Microsoft, and turns on whether training on copyrighted news content is lawful or instead infringes the publishers' rights. The claims at issue concern reproduction of protected works, removal or alteration of copyright management information, and related economic and moral rights, with fair use as the central defense in the United States and text-and-data-mining exceptions in the European Union.123

What this rests on5 independent sources
  • Evidence 18
  • Interpretation 2

In brief

  1. USA Today's publisher is among the news organizations suing OpenAI over the use of its journalism as AI training data, part of a wider wave of publisher-versus-AI copyright litigation.13

    Interpretation
  2. The claims turn on reproduction of protected works, copyright management information, and economic and moral rights.2

    Evidence-backed
  3. Fair use in the U.S. and the EU text-and-data-mining exception are the main defenses AI developers rely on.1

    Evidence-backed
  4. If owners prevail, developers could face billions in damages, injunctions, or orders to destroy models; if fair use succeeds, training on copyrighted data can continue.4

    Evidence-backed
  5. Licensing deals with major media organizations are already being struck alongside the litigation.1

    Evidence-backed

At a glance

The picture in numbers

Live · updated just now

Catalogued in one repository

41 controversies

41 controversies: ongoing legal controversies over copyright and data protection in AI training5

The evidence behind it

5 sources
  • Other studies and data5

When it was published

Newest from 2026

20242026
Sources on this page by kind and year
SourceKindYear
The Lawfulness of Using Inventions for Generative AI Training : A Case Study of a US Lawsuit against OpenAI and Perplexity AIOther studies and data2024
Umjetna inteligencija i mediji: The New York Times Company protiv kompanija OpenAI i MicrosoftOther studies and data2024
Assessing the feasibility of collective licensing of in-copyright works as training data for generative AI systems.Other studies and data2026
Legal frictions for data openness: Reflections from a case-study on re-use of the open web for AI trainingOther studies and data2025
Strength in Numbers: Group Copyright Registration of Online News Articles as a Weapon in the Battle Against AI InfringementOther studies and data2026

The community around it

Contributions
0
People
0
Following
0

Nobody has added anything yet. Experience, evidence or a different view would show up here.

What it means for you

Which fits you?

Pick the situation closest to yours. Each answer says what it rests on.

If you are a news publisher deciding whether to sue or license

copyright is the central legal weapon, but group registration of website updates (89 FR 311) can make registering and enforcing rights over online pages more practical, while deals with AI developers remain an alternative path.31

Evidence-backed

If you are an AI developer assessing training-data risk

the exposure is asymmetric: a successful fair use defense leaves existing models and further training free, while a loss could mean billions in damages, injunctions on further development, or destruction of models trained on infringing works.4

Evidence-backed

If you are a smaller developer, research centre or university building generative AI

any collective licensing mandate would apply to you too, not only to the large firms whose models are widely used today.4

Evidence-backed

If you are following how this litigation may resolve

the New York Times Company case against OpenAI and Microsoft, and parallel copyright-and-AI cases, are the closest guides to the direction the USA Today dispute may take.1

Evidence-backed

If you are a policymaker designing rules for AI training data

options discussed include requiring summaries of training datasets, defining developer and user responsibilities, AI limitation disclaimers, and non-exclusive blanket licensing through collective management organizations, synchronized with related policies.2

Evidence-backed

The full story · 4 chapters

01

What the lawsuit is about

AI summary:The case is one of several copyright disputes between news publishers and generative AI developers over using journalism to train large language models.

Evidence-backed

Evidence-backed: The case is part of a set of copyright disputes between news publishers and generative AI developers over whether journalism can be ingested and used to train large language models without permission or payment. The best-documented parallel is The New York Times Company's suit against OpenAI and Microsoft, which examines the origin and content of the dispute and looks at related cases on copyright and artificial intelligence for guidance on how such litigation may develop.1

Evidence-backed

Evidence-backed: Publishers treat copyright as their central weapon in this fight, but U.S. copyright procedures have lagged behind how news is published online. A new U.S. Copyright Office regulation, "Group Registrations of Updates to a News Website" (89 FR 311), lets a digital news publisher register a group of website updates by submitting a representative portion rather than the complete contents, which could make it more practical for outlets to register and then enforce rights over their pages while the larger AI-and-media legal questions remain unresolved.3

02

The claims at issue

AI summary:Publishers point to uncited sources, reproduced works and distorted works, with fair use in the U.S. and the EU mining exception as the main defenses.

Evidence-backed

Evidence-backed: Generative AI outputs can infringe copyright in several distinct ways: where the source is not cited, which violates copyright management information rules; where substantial portions of a work are reproduced, violating the rightholder's economic rights; or where a work is distorted in a way that harms the rightholder's honor, infringing moral rights. These are the categories of harm publishers point to when they argue that training and output generation cross legal lines.2

Evidence-backed

Evidence-backed: In the United States, the pivotal question is fair use, especially whether the use of works to train models is "transformative." In the European Union, the analogous question is whether the text-and-data-mining exception in Directive (EU) 2019/790 applies, an exception strongly backed by the text of the EU AI Regulation. These safe harbours are what AI developers rely on to argue that training on copyrighted news content is lawful.1

03

Training data and licensing

AI summary:The dispute spans many parallel controversies, contested collective licensing proposals, and deals already struck with major media organizations.

Evidence-backed

Evidence-backed: The training-data dispute has produced a large body of parallel controversies: one repository catalogues 41 ongoing legal controversies relating to copyright and data protection in the training of foundation AI models, and analyses how they affect or advance the openness of training datasets. The same work argues that legal strategies are needed both to limit data extractivism by well-resourced actors such as Big Tech and to enable community data sovereignty, and critically examines open, permissive and alternative licensing frameworks for training datasets.5

Evidence-backed

Evidence-backed: Collective licensing has been widely proposed as a compromise that would let generative AI systems be developed while compensating copyright owners, but its feasibility is contested. The normative, economic and practical problems are substantial, and a licensing mandate would affect not only the large firms whose models are widely used today but also start-ups, research centres, higher education developers and the general public.4

Evidence-backed

Evidence-backed: Alongside litigation, some AI developers have concluded agreements and even partnerships with large media organizations — for example Associated Press, Le Monde and News Corp — covering the use of their content in GPT-model training and technological cooperation. This shows that licensing deals are being struck even while the legal questions remain open.1

Evidence-backed

Evidence-backed: For jurisdictions looking to accommodate AI development while protecting rightholders, proposals include obligations for AI companies to release summaries of training datasets, rules defining the responsibilities of AI developers and users, disclaimers about AI's limitations, and a non-exclusive blanket licence through collective management organizations, synchronized with related policies to provide legal certainty.2

04

What could follow

AI summary:If fair use wins, training on copyrighted data continues; if owners win, developers could face damages, injunctions or orders to destroy models.

Evidence-backed

Evidence-backed: The stakes are large in both directions. If fair use defenses succeed, developers will be free to continue commercially exploiting models already built on copyrighted data and to use those data to train new models or fine-tune existing ones. If copyright owners prevail, developers may be liable for billions of dollars in damages, could be enjoined from further model development on in-copyright works, and could even be ordered to destroy models trained on infringing works.4

Ask this Sylo

Still wondering about something?

Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.

Behind this page

Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.

Discussion

Nobody has added anything yet. If you have experience, evidence or a different view, you could be the first.

Sources

Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.

  1. 1
    Umjetna inteligencija i mediji: The New York Times Company protiv kompanija OpenAI i Microsoft
    Research paper (Mešević)Published Oct 1, 2024Checked Oct 9, 2026
    “In addition, it addresses the possible safe harbours for AI-developers in this context. Firstly, in the European Union, in terms of the text and data mining exception from Directive (EU) 2019/790, which is strongly backed by the text of the EU AI Regulation. Secondly, in the United States of America, in terms of applying the legal doctrine of „fair use“, especially in the sense of the transformative use of works. The author elaborates in detail the circumstances of the origin and the content of the dispute and looks at parallel cases in the domain of relationship between copyright and artificial intelligence, which could provide some guidance regarding the direction of the outcome of the NYTC– case. In addition, the paper draws attention to positive examples of OpenAI concluding agreements and even creating partnerships with large media organizations (e.g. Associated Press, Le Monde or News Corp) regarding the use of their content in GPT-model training, as well as in the context of technological cooperation.”
  2. 2
    The Lawfulness of Using Inventions for Generative AI Training : A Case Study of a US Lawsuit against OpenAI and Perplexity AI
    JUSTISI (Ismantara & Silalahi)Published Dec 25, 2024Checked Oct 9, 2026
    “GAI outputs may also infringe copyright if: (1) the source is not cited, violating Article 7 on copyright management information; (2) substantial portions of the work are reproduced, violating the rightholder's economic rights under Article 9; or (3) the work is distorted in a way that harms the rightholder’s honor, infringing on moral rights under Article 5. To accommodate AI development, specific regulations integrating AI transparency principles outlined in SE Kominfo 9/2023 are required. These regulations could include obligations for AI companies to release summaries of training datasets, include Uni EropaLAs that define the responsibilities of AI developers and users, and provide disclaimers regarding AI's limitations. Regarding the fulfillment of rightholders’ economic rights, a non-exclusive blanket license through Collective Management Organizations (CMOs) as stipulated in Permenkumham 15/2024 is necessary. These regulations should be synchronized with related policies to establish legal certainty that adapts to technological advancements.”
  3. 3
    Strength in Numbers: Group Copyright Registration of Online News Articles as a Weapon in the Battle Against AI Infringement
    Utah law review (Farmer)Published Jan 1, 2026Checked Oct 9, 2026
    “In this fight, copyright protection remains an essential weapon for publishers to safeguard their content. However, U.S. copyright law has lagged behind the news industry’s integration of new technologies and mediums, with current procedures for staying abreast of copyright registration burdensome or infeasible. A new regulation from the U.S. Copyright Office, “Group Registrations of Updates to a News Website,” or 89 FR 311, allows a digital news publisher to collectively register a group of updates to a news website by submitting a representative portion of the content, rather than the complete contents of the website. It has the potential to resolve challenges standing in the way of online news publishers seeking this protection for their pages. I examine how published content is used by generative AI, current copyright protections, and how 89 FR 311 could strengthen copyright protections for news publishers while significant legal issues regarding generative artificial intelligence and the media await resolution.”
  4. 4
    Assessing the feasibility of collective licensing of in-copyright works as training data for generative AI systems.
    Proceedings of the National Academy of Sciences of the United States of America (Samuelson)Published Jul 20, 2026Checked Oct 9, 2026
    “If fair use defenses succeed, developers will be free to continue to commercially exploit models already built on copyrighted data as well as to use these data to train new models or fine-tune existing ones. If copyright owners prevail, developers may be liable for billions of dollars of damages. Developers could also be enjoined from further model development on in-copyright works and even ordered to destroy models trained on infringing works. Numerous commentators have proposed collective licensing as a compromise solution to the copyright-training-data dilemma. Other commentators have questioned the feasibility of such a compromise. This article discusses several proposals for collective licensing to enable development of generative AI systems while providing some compensation to copyright owners. It assesses the complex normative, economic, and practical problems that must be addressed if such a regime is to become feasible. It discusses the implications of a licensing mandate not only for the large firms whose models are widely used today, but also for start-ups, research centers, and higher education developers of generative AI systems, as well as the general public.”
  5. 5
    Legal frictions for data openness: Reflections from a case-study on re-use of the open web for AI training
    HAL (Le Centre pour la Communication Scientifique Directe) (Ramya)Published Mar 27, 2025Checked Oct 9, 2026
    “While techno-legal openness is necessary, this report argues that the political economy of data re-use also necessitates legal strategies that impose certain limits on data extractivism by well-resourced actors like Big Tech on the one hand, and enable community data sovereignty on the other hand. It contains a repository of 41 ongoing legal controversies relating to copyright and data protection related to training foundation AI models, together with a detailed analysis of how these legal controversies either impact or advance three-dimensional data openness of training datasets. It also contains a critical analysis of existing open licenses, permissive licenses, as well as certain alternative licensing frameworks for training datasets. While these licensing frameworks impose more obligations on re-users and necessitate more collective thinking on interoperability, these licensing frameworks together with other legal and institutional changes are nonetheless necessary for the creation of healthy digital and data commons, to realise the original promise of the open web as open for all.”

How it changed

Published 1 time since Oct 9, 2026.

  1. Version 2Oct 9, 2026Live now

    AI-prepared Starting Map from live research.

    • First published version.
Every version, side by side

Help improve it

The brief is open about what's uncertain. These are the specific gaps that new material would fill.

  • “What could follow” rests on one independent source

    A second, independent source that confirms or challenges it would make this part more reliable.

Open questions

  • What exactly does the USA Today complaint allege — which defendants, which works, which claims and what relief — beyond the general copyright framework described here?

    No answers yet

  • Will U.S. courts treat training on news content as transformative fair use, or find infringement?

    No answers yet

  • Will publishers increasingly settle through licensing deals like those with Associated Press, Le Monde and News Corp, or press litigation to judgment?

    No answers yet

  • How much will group registration of news website updates change publishers' ability to bring and sustain infringement claims?

    No answers yet

Around this topic

Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.

Add what you know

Sign in to add what you know. Reading stays open to everyone.

Ask this Sylo

Answers only from “Why is USA Today suing OpenAI over training data?”

Ask anything about this page. The AI reads only its reviewed brief, sources and contributions, cites what it used, and says when the page doesn't cover something.