Many generative AI systems use expressive works during development or deployment, placing copyright at the center of active private-law disputes. The question is whether the statement that “AI copyright litigation is the biggest active fight” is legally and factually defensible as of August 2, 2026. This analysis assumes “biggest” refers to the breadth of pending litigation, monetary and injunctive exposure, doctrinal uncertainty, and potential effects on model development and output design. The principal jurisdiction is United States federal copyright law. England and Wales and the European Union provide comparison.
Summary
United States
Copyright is the central active private-law fight over generative AI under the stated measures. No court has ranked it against privacy, competition, labor, safety, or public-law disputes.
No identified binding appellate decision as of the As-of Date decides whether training a modern general-purpose generative model on copyrighted expressive works is fair use. Section 107 requires use-specific balancing. Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508, 525–40 (2023).
Courts must separate acquisition, retention, training, retrieval, model distribution, outputs, and user conduct. Bartz treated training as fair use but rejected fair use for a permanent library built from pirated books. Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007, 1017–33 (N.D. Cal. 2025).
Market proof will decide many training cases. Meta prevailed against the named Kadrey plaintiffs because their record lacked meaningful market-harm evidence, not because all model training is lawful. Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026, 1036–60 (N.D. Cal. 2025).
Output claims remain independent of training claims. Detailed summaries, continuations, or near-verbatim responses may support infringement claims when they copy protected expression. In re OpenAI, Inc. Copyright Infringement Litigation, 2025 WL 3003339, at *7–10 (S.D.N.Y. Oct. 27, 2025).
Cox tightened contributory-liability doctrine. A service provider must intend infringement through inducement or a service tailored to infringement. Knowledge and inaction alone do not suffice. Cox Communications, Inc. v. Sony Music Entertainment, 607 U.S. 583, 592–96 (2026).
Remedies create exceptional settlement pressure. Bartz ended with final approval of a $1.5 billion non-reversionary settlement covering 482,460 listed works, plus specified file destruction. The settlement is not a merits precedent. Order Granting Final Approval, Bartz, No. 24-cv-05417-AMO, Dkt. 680, at 3–5, 10–11, 16–17, 22–23 (N.D. Cal. July 20, 2026).
England and Wales; European Union
Territorial rules diverge. The United Kingdom’s text-and-data-mining exception is narrow. EU law permits defined mining uses subject to lawful access and rights reservations, while Article 53 of the AI Act adds copyright-policy and training-content-summary duties.
United States; England and Wales; European Union
The decisive records are source manifests, licenses, acquisition dates, retained copies, rights reservations, training configurations, memorization tests, retrieval settings, output exemplars, marketing, notices, and registration data.
Analysis by Issue
Why copyright is the central active AI fight
Copyright is the best candidate for the largest active private-law dispute over generative AI. That assessment rests on four measurable features.
First, verified disputes span books, legal publishing, software code, and photography. Bartz, 787 F. Supp. 3d 1007; Ross, 765 F. Supp. 3d 382; Doe 1 v. GitHub, Inc., No. 24-7700; Getty Images (US), Inc. v. Stability AI Ltd [2025] EWHC 2863 (Ch). The OpenAI MDL also covers novels, news articles, and video transcripts. Transfer Order at 1–3, In re OpenAI, Inc. Copyright Infringement Litigation, MDL No. 3143, Dkt. 85 (J.P.M.L. Apr. 3, 2025). A single model may therefore face claims from several rightsholder groups.
Second, the Copyright Act combines aggregate scale with work-by-work remedies. A claimant may seek actual damages and attributable profits. An eligible claimant may instead elect statutory damages for each infringed work. 17 U.S.C. § 504(a)–(c). Courts may also grant injunctions, impoundment, disposition, costs, and attorney’s fees. 17 U.S.C. §§ 502–505.
Third, the controlling appellate rule for modern generative training remains unsettled. District courts have issued materially different rulings on different records. The Third Circuit has heard the Ross appeal, but no opinion had issued by the as-of date. No identified Ninth Circuit merits decision had established a general training rule.
Fourth, copyright decisions can alter data acquisition, dataset retention, model design, retrieval, filtering, licensing, and product economics. Those consequences reach the construction and operation of the system.
The superlative still needs limits. Privacy enforcement may carry larger public penalties in a given jurisdiction. Competition cases may reshape market structure. Safety rules may impose broader public duties. Labor disputes may affect more workers. “Biggest” therefore states an evaluative thesis, not an adjudicated fact.
The more exact proposition is this: copyright is the central active private-law fight because it tests whether developers may copy, retain, transform, and deploy protected expression at industrial scale.
The dispute covers separate acts
A court should not decide “AI training” as one undivided act. Section 106 assigns separate rights, and Section 107 evaluates the challenged use. 17 U.S.C. §§ 106–107. Warhol directs courts to examine the specific allegedly infringing use. The same copying may be fair for one purpose and unfair for another. 598 U.S. at 533–40.
A complete act map includes acquisition, download, scanning, format conversion, deduplication, retention, training, fine-tuning, checkpoint creation, retrieval-augmented generation, model-weight distribution, generated outputs, and user republication. Each act may involve different copies, actors, purposes, defenses, and markets.
Bartz demonstrates the point. The court held that Anthropic’s use of books to train Claude was fair use on the summary-judgment record. It also treated scanning purchased books as fair format conversion under stated controls. It denied fair-use protection to permanent library copies obtained from pirate repositories. 787 F. Supp. 3d at 1017–33.
Kadrey used a different analytical path. The court entered summary judgment for Meta on the named plaintiffs’ training claim. Its March 25, 2026 order later stressed that Meta won because the plaintiffs lacked meaningful market-harm evidence. The court did not announce a rule that training protected works is always fair. Order Granting Motion for Leave to File Fourth Amended Complaint, Kadrey, No. 3:23-cv-03417-VC, Dkt. 700, at 3–4 (N.D. Cal. Mar. 25, 2026).
The separation question remains contested. Bartz treated pirated acquisition and library retention as distinct from training. Kadrey assessed alleged BitTorrent acquisition within a different record and kept distribution and contributory claims alive. Neither district ruling binds another court on a general rule.
This act-by-act method has a concrete consequence. A developer may prevail on training and still face liability for source acquisition, permanent retention, output copying, distribution, or inducement. A rightsholder may lose a broad training theory and still prevail on a narrower act.
Training fair use has no universal answer
Section 107 identifies four factors: purpose and character, nature of the copyrighted work, amount and substantiality, and market effect. The inquiry is equitable and use-specific. 17 U.S.C. § 107; Google LLC v. Oracle America, Inc., 593 U.S. 1, 18–21, 32–40 (2021).
Warhol limits broad claims of transformation. A new meaning or technical function does not end the first-factor inquiry. Courts compare the specific purposes, commercial character, justification for copying, and substitution risk. Shared commercial purposes can weigh against fair use. 598 U.S. at 525–40.
Bartz supplies a favorable merits ruling for modern language-model training. The court found a highly different training purpose and no allegation that Claude supplied infringing copies to users. It treated the training copies as inputs used to create a model capable of producing new text. 787 F. Supp. 3d at 1017–24.
Bartz does not create a safe harbor. It is one district ruling on one record. Its favorable training holding coexisted with an adverse ruling on pirated library copies. The district court’s final-approval order resolved the remaining class claims without appellate review of the merits order.
Kadrey gives developers a second favorable result, but its limiting language is material. The named plaintiffs failed to prove cognizable market harm on their summary-judgment record. The court later stated that this evidentiary failure, rather than a categorical rule, produced Meta’s victory. Kadrey, Dkt. 700, at 3–4.
Ross points the other way on a narrower technology. Ross used Bulk Memos derived from Westlaw headnotes to develop a competing legal-search product. The court held that the commercial, nontransformative use and market effect defeated fair use. It expressly limited its discussion to non-generative AI. Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc., 765 F. Supp. 3d 382, 398–401 (D. Del. 2025), interlocutory appeal pending, No. 25-2153 (3d Cir. argued June 11, 2026).
Ross matters because it links purpose to competitive substitution. Its facts differ from general-purpose model training. The product competed directly with Westlaw, and the copied materials served a closely related function. A court should not transpose that result without comparing the works, product, use, and market.
The current rule is therefore conditional. Training has a stronger fair-use case when it serves a distinct analytical function, uses no more than needed, controls memorization, and does not substitute for protected works. Separate liability risk remains when the developer acquires copies unlawfully, retains them for unrelated uses, or releases substitutive outputs.
Data provenance may decide liability first
Source provenance is not a fifth fair-use factor. It can decide a separate infringement claim and supply facts about purpose, justification, and market effect.
A developer may obtain the same work from a licensed database, a purchased copy, an open repository, a scraped website, or a pirate archive. Those routes create different rights, contractual limits, authenticity concerns, and proof records. “Publicly accessible” does not mean “licensed” or “free of copyright.”
Bartz treated the source route as legally consequential. The court rejected the claim that later training automatically excused earlier piracy and permanent retention. It required a justification for each use. 787 F. Supp. 3d at 1024–33.
The Bartz settlement shows the practical leverage associated with provenance evidence on this record. The court approved a $1.5 billion fund and destruction of specified LibGen and PiLiMi files and copies derived from them, subject to preservation duties. The release covered defined past input claims. It excluded output claims and future conduct after August 25, 2025. Bartz, Dkt. 680, at 3–5, 10–11, 16–17.
The settlement does not prove infringement. Rule 23 approval tests whether a settlement is fair, reasonable, and adequate. It does not adjudicate every released claim. The amount still shows the practical exposure created by a documented source list, registered works, class treatment, and unresolved statutory damages.
A defensible provenance record should identify the work, source, acquisition date, license, territorial scope, restrictions, version, hash, retention path, training runs, deletion history, and rights reservations. Missing records can prevent a developer from proving lawful access, licensed use, or factual distinctions needed for fair use.
Market proof will decide many cases
The fourth factor asks about the effect of the challenged use upon the potential market for or value of the copyrighted work. 17 U.S.C. § 107(4). Warhol also treats substitution risk as relevant to purpose and justification. 598 U.S. at 531–40.
Kadrey makes market proof central. The plaintiffs argued that model training impaired licensing markets and enabled a flood of competing works. The court found their proof insufficient on the named plaintiffs’ record. It did not foreclose a better-supported market theory. 788 F. Supp. 3d at 1036–60; Kadrey, Dkt. 700, at 3–4.
Ross presents a clearer substitution record. Ross built a legal-search competitor using materials derived from Westlaw headnotes. The first and fourth factors favored Thomson Reuters. 765 F. Supp. 3d at 398–401.
A rightsholder needs evidence tied to the challenged system and work. Useful proof includes lost subscriptions, displaced sales, price effects, failed licenses, comparable licenses, customer switching, retrieval frequency, near-verbatim outputs, memorization rates, and causation analysis. A general claim that AI creates more competition is weaker.
A developer needs equally concrete proof. Relevant material includes output testing, low retrieval rates, anti-memorization measures, product differentiation, nonexpressive functions, licensed datasets, and evidence that users still obtain the original works. The inquiry concerns copyright markets, not protection from lawful competition.
An asserted training-license market cannot automatically control the fourth factor. Treating every requested license as proof of market harm would make fair use circular. Kadrey, No. 23-cv-03417-VC, Dkt. 598, at 26–27 (N.D. Cal. June 25, 2025). A claimant should prove a cognizable market and likely injury with actual transactions, customary licensing, or substitution evidence.
Outputs create a separate infringement track
A fair-use ruling for training does not immunize model outputs. Output liability turns on ownership, copying, protected expression, substantial similarity, defenses, and the actor responsible for the challenged output.
The OpenAI multidistrict litigation illustrates the pleading threshold. The district court held that allegations and submitted outputs could support actual copying and substantial similarity. Some outputs included detailed summaries and sequel-style material tied to protected characters and plot expression. The court denied dismissal of the output-based direct claim. It made no fair-use finding. In re OpenAI, 2025 WL 3003339, at *7–10.
That ruling is not a merits judgment. It asks whether a reasonable jury could find substantial similarity on the pleaded materials. OpenAI may contest copying, protectability, similarity, fair use, causation, and responsibility after discovery.
Copyright does not protect an idea, procedure, process, system, method of operation, concept, principle, or discovery. 17 U.S.C. § 102(b). A style-only resemblance may therefore fail unless the output also copies protectable expression. A response may infringe when it reproduces protected text, characters, selection, arrangement, or other expression. Retrieval and memorization can increase that risk.
Risk rises when a system returns long passages, detailed substitutes, continuations using protected expression, paywalled articles, or images closely matching training works. Risk falls when outputs remain remote from protectable expression and serve a distinct function.
Product design affects proof. Retrieval corpora, chunk size, similarity thresholds, system prompts, decoding settings, filters, and citation features can determine whether the service supplies expressive substitutes. Logs can show whether an output was ordinary, adversarially elicited, repeated, blocked, or encouraged.
Secondary liability narrowed in 2026
Cox now supplies the controlling national rule for contributory copyright liability. A provider is liable for user infringement only when it intended its service to be used for infringement. Intent may appear through active inducement or a service tailored to infringement. Mere knowledge and failure to take preventive steps are insufficient. Cox, 607 U.S. at 592–96.
Cox did not concern generative AI. Its rule still applies when rightsholders claim that a model provider contributed to users’ direct infringement. The claimant must first establish predicate infringement by users.
General-purpose models have many lawful uses. That fact supports the defense that the service is not tailored to infringement. A provider also benefits from marketing and design that do not encourage copying.
The risk changes when evidence shows active promotion of infringing uses. Prompt libraries aimed at reproducing protected works, advertising centered on imitation, repeated facilitation of known copying, or features designed for unauthorized substitution may support intent.
Content filters, notices, and repeat-user controls remain relevant evidence. Cox means that inadequate action alone does not establish contributory intent. A failed filter is not equivalent to inducement.
Direct liability remains separate. A provider that itself makes an infringing copy does not receive Cox’s intent protection for that act. The classification of model generation, retrieval, and delivery therefore matters.
Section 1202 remains a narrower front
Section 1202 prohibits specified acts involving false, removed, or altered copyright management information. A civil claim under Section 1202(b) requires knowledge of the removal or alteration and knowledge that distribution will induce, enable, facilitate, or conceal infringement. 17 U.S.C. § 1202(b).
This double-scienter requirement makes the claim narrower than direct infringement. A plaintiff must identify qualifying information, the prohibited act, the defendant’s state of mind, and the connection to infringement.
The pending Doe v. GitHub appeal may clarify how these requirements apply to code and AI-assisted outputs. Doe 1 v. GitHub, Inc., No. 24-7700 (9th Cir. argued Feb. 11, 2026). No appellate opinion had issued by the as-of date.
Section 1203 permits actual damages and profits or statutory damages within stated ranges for Section 1202 violations. The possible unit of violation can create substantial exposure. Courts still require proof of the statutory elements.
Remedies create settlement pressure
The Copyright Act authorizes injunctions, impoundment and disposition, actual damages and profits, statutory damages, costs, and discretionary attorney’s fees. 17 U.S.C. §§ 502–505.
Statutory damages generally range from $750 to $30,000 for each infringed work. A court may increase the maximum to $150,000 for willful infringement. It may reduce the amount to $200 for innocent infringement. 17 U.S.C. § 504(c). The unit is the infringed work, not each copy or training pass.
Section 412 can bar statutory damages and attorney’s fees when registration timing does not satisfy the statute. 17 U.S.C. § 412. Registration, ownership chains, group registrations, and definitive work lists can therefore determine case value before fair use is reached.
Bartz gives the clearest measure of settlement scale. The final order approved a $1.5 billion non-reversionary fund for 482,460 listed works. It estimated about $3,000 per work before costs and fees and required specified file destruction. Bartz, Dkt. 680, at 3–5, 16–17.
That result does not set a damages tariff. It reflected defined works, registration criteria, class certification, litigation risk, and a negotiated release. Other cases may fail on ownership, registration, causation, fair use, class treatment, or proof of copying.
Section 502 makes injunctions discretionary. 17 U.S.C. § 502(a). A court could target data copies, retrieval, outputs, or distribution without ordering destruction of an entire model. None of the United States decisions reviewed for this analysis requires automatic model deletion solely because training data were unlawfully copied.
Human authorship is related but less unsettled
The D.C. Circuit held that a non-human machine cannot be the sole author under the Copyright Act. Thaler v. Perlmutter, 130 F.4th 1039, 1045–54 (D.C. Cir. 2025), cert. denied, No. 25-449 (U.S. Mar. 2, 2026).
Thaler does not bar protection for every AI-assisted work. Copyright may cover human-authored selection, arrangement, revision, or expression when the record identifies those contributions. The Copyright Office and courts must examine the claimed human authorship.
This issue affects ownership and enforcement, but it is narrower than training and output liability. The rule for a work attributed solely to a machine is now comparatively clear. Disputes over human control and contribution remain fact-specific.
Cross-border rules diverge
England and Wales do not provide a broad commercial mining exception. Section 29A of the Copyright, Designs and Patents Act 1988 covers copies made for computational analysis under lawful access, for the sole purpose of noncommercial research, subject to statutory conditions. Copyright, Designs and Patents Act 1988, c. 48, § 29A.
Getty Images v. Stability AI did not decide whether model training in the United Kingdom was lawful. Getty abandoned the training claim because the record did not establish United Kingdom training. It also abandoned its output copyright claim. The High Court held, on the remaining secondary-infringement theory, that the final model weights were not infringing copies because they did not store or reproduce the protected works. Getty Images (US), Inc. v. Stability AI Ltd [2025] EWHC 2863 (Ch) [599]–[604], [752]–[754].
That holding concerns the statutory meaning of an infringing copy on the trial record. It does not license the copying used to train a model. It also does not establish a universal technical fact about every set of model weights.
EU law uses defined text-and-data-mining exceptions. Article 3 of Directive 2019/790 covers qualifying research organizations and cultural heritage institutions conducting scientific research with lawful access. Article 4 covers broader mining of lawfully accessible works when rightsholders have not expressly reserved their rights in an appropriate manner. Directive (EU) 2019/790, arts. 3–4.
Article 53 of Regulation 2024/1689 requires providers of general-purpose AI models to maintain a policy for compliance with Union copyright law. That policy must address Article 4(3) rights reservations. Providers must also publish a sufficiently detailed summary of training content. Regulation (EU) 2024/1689, art. 53.
The EU AI Act duties applied to Chapter V from August 2, 2025. The Regulation’s general application date is August 2, 2026. These duties add compliance and disclosure obligations. They do not create a blanket copyright license.
A developer therefore cannot apply one global answer. Lawful access, reservations, research status, place of copying, model distribution, output location, and national implementation can change the result.
Likely direction and operating consequences
United States courts will probably continue to decide separate acts on product-specific records. Warhol requires use-specific analysis. Bartz, Kadrey, Ross, and the OpenAI output order point to the same procedural reality despite different outcomes.
The next major appellate decisions may come from Ross and Doe. Ross may address fair use in a non-generative training dispute. Doe may clarify Section 1202. Neither appeal necessarily resolves general-purpose generative training across media.
Cox already changes output and user-liability claims. Claimants must prove inducement or design tailored to infringement for contributory liability. Developers must still answer direct-copying claims and any evidence of active encouragement.
Developers should maintain work-level provenance, honor applicable rights reservations, test memorization and retrieval, limit substitutive outputs, document license scope, control marketing, and preserve relevant records. Those acts affect both liability and proof.
Rightsholders should register promptly, document ownership, identify definitive works, preserve output exemplars, and develop product-specific market evidence. General objections to AI use will not replace proof of copying, protectable expression, market harm, and remedy eligibility.
The statement under review is therefore supportable with one correction. AI copyright litigation is not proven to be the largest legal dispute by every measure. It is the largest active private-law fight over the economic inputs and expressive outputs of generative AI under the stated measures.
