Why is USA Today suing OpenAI over training data?
News publishers are suing OpenAI over using their journalism to train AI, and the outcome could decide whether that training is legal.
Covers: Covers the USA Today lawsuit against OpenAI, including the publishers involved, the copyright and related claims, the training data at issue, and the broader context of news publishers suing AI companies. Does not provide legal advice or predict the final outcome of the case.
3 free full reads left this month. Join or upgrade
The short answer
Interpretation AI-prepared starting mapUSA Today's publisher is among the news organizations suing OpenAI over the use of its journalism as training data for generative AI models. The dispute sits inside a broader wave of copyright litigation between publishers and AI developers, including The New York Times Company's case against OpenAI and Microsoft, and turns on whether training on copyrighted news content is lawful or instead infringes the publishers' rights. The claims at issue concern reproduction of protected works, removal or alteration of copyright management information, and related economic and moral rights, with fair use as the central defense in the United States and text-and-data-mining exceptions in the European Union.123
- Evidence 18
- Interpretation 2
In brief
The claims turn on reproduction of protected works, copyright management information, and economic and moral rights.2
Evidence-backedFair use in the U.S. and the EU text-and-data-mining exception are the main defenses AI developers rely on.1
Evidence-backedIf owners prevail, developers could face billions in damages, injunctions, or orders to destroy models; if fair use succeeds, training on copyrighted data can continue.4
Evidence-backedLicensing deals with major media organizations are already being struck alongside the litigation.1
Evidence-backed
At a glance
The picture in numbers
Live · updated just now
41 controversies
The evidence behind it
5 sources- Other studies and data5
When it was published
Newest from 2026
| Source | Kind | Year |
|---|---|---|
| The Lawfulness of Using Inventions for Generative AI Training : A Case Study of a US Lawsuit against OpenAI and Perplexity AI | Other studies and data | 2024 |
| Umjetna inteligencija i mediji: The New York Times Company protiv kompanija OpenAI i Microsoft | Other studies and data | 2024 |
| Assessing the feasibility of collective licensing of in-copyright works as training data for generative AI systems. | Other studies and data | 2026 |
| Legal frictions for data openness: Reflections from a case-study on re-use of the open web for AI training | Other studies and data | 2025 |
| Strength in Numbers: Group Copyright Registration of Online News Articles as a Weapon in the Battle Against AI Infringement | Other studies and data | 2026 |
The community around it
- Contributions
- 0
- People
- 0
- Following
- 0
Nobody has added anything yet. Experience, evidence or a different view would show up here.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you are a news publisher deciding whether to sue or license
copyright is the central legal weapon, but group registration of website updates (89 FR 311) can make registering and enforcing rights over online pages more practical, while deals with AI developers remain an alternative path.31
Evidence-backedIf you are an AI developer assessing training-data risk
the exposure is asymmetric: a successful fair use defense leaves existing models and further training free, while a loss could mean billions in damages, injunctions on further development, or destruction of models trained on infringing works.4
Evidence-backedIf you are a smaller developer, research centre or university building generative AI
any collective licensing mandate would apply to you too, not only to the large firms whose models are widely used today.4
Evidence-backedIf you are following how this litigation may resolve
the New York Times Company case against OpenAI and Microsoft, and parallel copyright-and-AI cases, are the closest guides to the direction the USA Today dispute may take.1
Evidence-backedIf you are a policymaker designing rules for AI training data
options discussed include requiring summaries of training datasets, defining developer and user responsibilities, AI limitation disclaimers, and non-exclusive blanket licensing through collective management organizations, synchronized with related policies.2
Evidence-backedThe full story · 4 chapters
01
What the lawsuit is about
AI summary:The case is one of several copyright disputes between news publishers and generative AI developers over using journalism to train large language models.
Evidence-backed: The case is part of a set of copyright disputes between news publishers and generative AI developers over whether journalism can be ingested and used to train large language models without permission or payment. The best-documented parallel is The New York Times Company's suit against OpenAI and Microsoft, which examines the origin and content of the dispute and looks at related cases on copyright and artificial intelligence for guidance on how such litigation may develop.1
Evidence-backed: Publishers treat copyright as their central weapon in this fight, but U.S. copyright procedures have lagged behind how news is published online. A new U.S. Copyright Office regulation, "Group Registrations of Updates to a News Website" (89 FR 311), lets a digital news publisher register a group of website updates by submitting a representative portion rather than the complete contents, which could make it more practical for outlets to register and then enforce rights over their pages while the larger AI-and-media legal questions remain unresolved.3
02
The claims at issue
AI summary:Publishers point to uncited sources, reproduced works and distorted works, with fair use in the U.S. and the EU mining exception as the main defenses.
Evidence-backed: Generative AI outputs can infringe copyright in several distinct ways: where the source is not cited, which violates copyright management information rules; where substantial portions of a work are reproduced, violating the rightholder's economic rights; or where a work is distorted in a way that harms the rightholder's honor, infringing moral rights. These are the categories of harm publishers point to when they argue that training and output generation cross legal lines.2
Evidence-backed: In the United States, the pivotal question is fair use, especially whether the use of works to train models is "transformative." In the European Union, the analogous question is whether the text-and-data-mining exception in Directive (EU) 2019/790 applies, an exception strongly backed by the text of the EU AI Regulation. These safe harbours are what AI developers rely on to argue that training on copyrighted news content is lawful.1
03
Training data and licensing
AI summary:The dispute spans many parallel controversies, contested collective licensing proposals, and deals already struck with major media organizations.
Evidence-backed: The training-data dispute has produced a large body of parallel controversies: one repository catalogues 41 ongoing legal controversies relating to copyright and data protection in the training of foundation AI models, and analyses how they affect or advance the openness of training datasets. The same work argues that legal strategies are needed both to limit data extractivism by well-resourced actors such as Big Tech and to enable community data sovereignty, and critically examines open, permissive and alternative licensing frameworks for training datasets.5
Evidence-backed: Collective licensing has been widely proposed as a compromise that would let generative AI systems be developed while compensating copyright owners, but its feasibility is contested. The normative, economic and practical problems are substantial, and a licensing mandate would affect not only the large firms whose models are widely used today but also start-ups, research centres, higher education developers and the general public.4
Evidence-backed: Alongside litigation, some AI developers have concluded agreements and even partnerships with large media organizations — for example Associated Press, Le Monde and News Corp — covering the use of their content in GPT-model training and technological cooperation. This shows that licensing deals are being struck even while the legal questions remain open.1
Evidence-backed: For jurisdictions looking to accommodate AI development while protecting rightholders, proposals include obligations for AI companies to release summaries of training datasets, rules defining the responsibilities of AI developers and users, disclaimers about AI's limitations, and a non-exclusive blanket licence through collective management organizations, synchronized with related policies to provide legal certainty.2
04
What could follow
AI summary:If fair use wins, training on copyrighted data continues; if owners win, developers could face damages, injunctions or orders to destroy models.
Evidence-backed: The stakes are large in both directions. If fair use defenses succeed, developers will be free to continue commercially exploiting models already built on copyrighted data and to use those data to train new models or fine-tune existing ones. If copyright owners prevail, developers may be liable for billions of dollars in damages, could be enjoined from further model development on in-copyright works, and could even be ordered to destroy models trained on infringing works.4
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1Umjetna inteligencija i mediji: The New York Times Company protiv kompanija OpenAI i MicrosoftResearch paper (Mešević)Published Oct 1, 2024Checked Oct 9, 2026
“In addition, it addresses the possible safe harbours for AI-developers in this context. Firstly, in the European Union, in terms of the text and data mining exception from Directive (EU) 2019/790, which is strongly backed by the text of the EU AI Regulation. Secondly, in the United States of America, in terms of applying the legal doctrine of „fair use“, especially in the sense of the transformative use of works. The author elaborates in detail the circumstances of the origin and the content of the dispute and looks at parallel cases in the domain of relationship between copyright and artificial intelligence, which could provide some guidance regarding the direction of the outcome of the NYTC– case. In addition, the paper draws attention to positive examples of OpenAI concluding agreements and even creating partnerships with large media organizations (e.g. Associated Press, Le Monde or News Corp) regarding the use of their content in GPT-model training, as well as in the context of technological cooperation.”
- 2The Lawfulness of Using Inventions for Generative AI Training : A Case Study of a US Lawsuit against OpenAI and Perplexity AIJUSTISI (Ismantara & Silalahi)Published Dec 25, 2024Checked Oct 9, 2026
“GAI outputs may also infringe copyright if: (1) the source is not cited, violating Article 7 on copyright management information; (2) substantial portions of the work are reproduced, violating the rightholder's economic rights under Article 9; or (3) the work is distorted in a way that harms the rightholder’s honor, infringing on moral rights under Article 5. To accommodate AI development, specific regulations integrating AI transparency principles outlined in SE Kominfo 9/2023 are required. These regulations could include obligations for AI companies to release summaries of training datasets, include Uni EropaLAs that define the responsibilities of AI developers and users, and provide disclaimers regarding AI's limitations. Regarding the fulfillment of rightholders’ economic rights, a non-exclusive blanket license through Collective Management Organizations (CMOs) as stipulated in Permenkumham 15/2024 is necessary. These regulations should be synchronized with related policies to establish legal certainty that adapts to technological advancements.”
- 3Strength in Numbers: Group Copyright Registration of Online News Articles as a Weapon in the Battle Against AI InfringementUtah law review (Farmer)Published Jan 1, 2026Checked Oct 9, 2026
“In this fight, copyright protection remains an essential weapon for publishers to safeguard their content. However, U.S. copyright law has lagged behind the news industry’s integration of new technologies and mediums, with current procedures for staying abreast of copyright registration burdensome or infeasible. A new regulation from the U.S. Copyright Office, “Group Registrations of Updates to a News Website,” or 89 FR 311, allows a digital news publisher to collectively register a group of updates to a news website by submitting a representative portion of the content, rather than the complete contents of the website. It has the potential to resolve challenges standing in the way of online news publishers seeking this protection for their pages. I examine how published content is used by generative AI, current copyright protections, and how 89 FR 311 could strengthen copyright protections for news publishers while significant legal issues regarding generative artificial intelligence and the media await resolution.”
- 4Assessing the feasibility of collective licensing of in-copyright works as training data for generative AI systems.Proceedings of the National Academy of Sciences of the United States of America (Samuelson)Published Jul 20, 2026Checked Oct 9, 2026
“If fair use defenses succeed, developers will be free to continue to commercially exploit models already built on copyrighted data as well as to use these data to train new models or fine-tune existing ones. If copyright owners prevail, developers may be liable for billions of dollars of damages. Developers could also be enjoined from further model development on in-copyright works and even ordered to destroy models trained on infringing works. Numerous commentators have proposed collective licensing as a compromise solution to the copyright-training-data dilemma. Other commentators have questioned the feasibility of such a compromise. This article discusses several proposals for collective licensing to enable development of generative AI systems while providing some compensation to copyright owners. It assesses the complex normative, economic, and practical problems that must be addressed if such a regime is to become feasible. It discusses the implications of a licensing mandate not only for the large firms whose models are widely used today, but also for start-ups, research centers, and higher education developers of generative AI systems, as well as the general public.”
- 5Legal frictions for data openness: Reflections from a case-study on re-use of the open web for AI trainingHAL (Le Centre pour la Communication Scientifique Directe) (Ramya)Published Mar 27, 2025Checked Oct 9, 2026
“While techno-legal openness is necessary, this report argues that the political economy of data re-use also necessitates legal strategies that impose certain limits on data extractivism by well-resourced actors like Big Tech on the one hand, and enable community data sovereignty on the other hand. It contains a repository of 41 ongoing legal controversies relating to copyright and data protection related to training foundation AI models, together with a detailed analysis of how these legal controversies either impact or advance three-dimensional data openness of training datasets. It also contains a critical analysis of existing open licenses, permissive licenses, as well as certain alternative licensing frameworks for training datasets. While these licensing frameworks impose more obligations on re-users and necessitate more collective thinking on interoperability, these licensing frameworks together with other legal and institutional changes are nonetheless necessary for the creation of healthy digital and data commons, to realise the original promise of the open web as open for all.”
How it changed
Published 1 time since Oct 9, 2026.
- Version 2Oct 9, 2026Live now
AI-prepared Starting Map from live research.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
“What could follow” rests on one independent source
A second, independent source that confirms or challenges it would make this part more reliable.
Open questions
What exactly does the USA Today complaint allege — which defendants, which works, which claims and what relief — beyond the general copyright framework described here?
No answers yet
Will U.S. courts treat training on news content as transformative fair use, or find infringement?
No answers yet
Will publishers increasingly settle through licensing deals like those with Associated Press, Le Monde and News Corp, or press litigation to judgment?
No answers yet
How much will group registration of news website updates change publishers' ability to bring and sustain infringement claims?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.