Introduction
The emergence of Large Language Models (LLMs) has inevitably prompted us to reconsider the applicability of Copyright law to Artificial Intelligence (“AI”). Most text-based AI models, such as ChatGPT Claude have been trained on large amounts of text, majority of which is copyrighted. So, there are questions concerning the shielding of Intellectual Property Rights (“IPR”) and rights to use data ethically.
On 24th June 2025, the United States District Court gave its decision in Bartz v. Anthropic PBC. The Court, for the first time, has taken a substantive approach to copyright infringement while the legal nuances of artificial intelligence training are still being explored. In a finely-balanced decision, Judge William Alsup ruled that while using copyrighted works to train LLMs may potentially qualify as fair use, the Court denied summary judgment where pirated copies were involved. In this case, the training set that Anthropic made consisted of more than 7 million illegitimate copies of e-books supplied by the black-market information libraries This inevitably attracted the attention of the court to digital piracy. Though the court was not opposed to considering fair use regarding legitimately obtained material, it denied granting summary judgment on the use of pirated content, on the ground of illegal acquisition.
This article analyses the legal implications of the Anthropic ruling by examining relevant legal implications of the Anthropic ruling. It focuses on, firstly, whether the utilization of the material under the copyright to train the Large Language Models (LLMs) can be described as fair usage, secondly, whether the courts can or should disqualify claims of fair use due to the acts of piracy, thirdly, whether there is a possibility of Indian courts aligning with this judgement and lastly, the concise evaluation of Fair use doctrine and the act of piracy concerning such cases.
Fair Use Doctrine: Legal Structure
The Fair Use doctrine, as contended in the Anthropic case, grants an individual the right to make a limited copy of a copyrighted work without the permission of the holder, provided that it will be used within the limits of the provision. This four-factor doctrine incorporates the purpose and character of your use, the nature of the copyrighted work, the amount and substantiality of the portion taken, and the effect of the use upon the potential market.
In the setting of AI training, such factors are complicated. The LLMs do not generate the content verbatim; rather, they are trained to understand patterns and churn out novel outputs. The major legal issue is whether such computational reading of the protected works can be characterised as a transformative use when the model generated does not present the underlying works in its copyrighted form.
In the case of Sony v. Universal City Studios, the Supreme Court held personal home video recording to be fair use, and in the case of Authors Guild v. Google, it was held that digitizing books so as to provide a searchable index was transformative, and it did not hurt the market. Collectively, these cases show how the courts use the doctrine to promote technological innovation that does not affect the worth of original works in a negative way.
The evolving copyright jurisprudence in the USA has made machine processing a form of transformative analysis. . However, Jane C. Ginsburg in her article noted that “a finding of ‘transformativeness’ often foreordained the ultimate outcome.” she has warned against too expansive a reading of “transformative”. Furthermore, while making a case brief on Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith, Pamela Samuelson said that, “Reaffirming transformativeness as a significant consideration in fair use case law would be very welcome.”
Piracy and the Fair Use Calculus: A Doctrinal Shift
The fair use debate in Bartz v. Anthropic was complicated by an additional, critical factor: digital piracy, which directly undermined the legitimacy of the AI training process. Digital piracy refers to the illegal copying or distribution of copyrighted material via the Internet.
The U.S. District Court noted that using copyrighted works by Anthropic for AI models training could have been considered defensible if it was lawfully procured. But the fact that Anthropic was using more than 7 million pirated e-books on shadow libraries such as Library Genesis and Z library weakened its fair use defence. This raises a question of, whether the courts can or should disqualify claims of fair use due to the acts of piracy?
In Harper & Row v. Nation Enterprises, the Supreme Court held that the unauthorised publication of President Ford’s memoir, although newsworthy, could not be considered fair use because it was maliciously procured. The implication of piracy is a conscious disregard of authors’ rights, and the incentive created by permitting a defendant to derive benefit from such a practice would be excessive and dangerous, especially where training sets are measured in the millions.
The court may disregard the fair use defences based on pirate activities, and the disqualifications, therefore, are not only permissible but also doctrinally as well as normatively sound. Courts must ensure that fair use is not the tool used to launder unlawful acquisition, particularly in large-stakes, commercially inclined technology that effectively rides on the intellectual effort of others.
Cross-border Compliance
Although the U.S. District Court did not overturn the statement that training on lawfully acquired copyrighted materials may be considered fair use because the non-expressive character of them poses a transformative effect, it did state that the incorporation of pirated content, especially that of sites, undermines any plausible fair use defence. The compliance requirement is not restricted to the United States. India’s legal threshold is even narrower, governed by the fair dealing exceptions under the Copyright Act, 1957. Limited exceptions to copyright under the Act are generally limited to personal use, research and education, as opposed to commercial AI training. Historically, the courts in India have been quite conservative in applying exceptions to copyright. Were an Anthropic case to appear before an Indian court, the consequences would have been no different. In the case of Super Cassettes v. MySpace, the Indian judiciary emphasised the responsibility of internet sites and the strict regulation of unauthorised use of digital content. The use of pirated datasets, irrespective of the ultimate purpose of the model, is more likely to be deemed insufficient to allow a fair-dealing defence.
What Bartz indicates, then, is the shared expectation around the world. AI developers need to exercise due diligence when choosing the source of content. The legal regime after Bartz requires establishing the institutionalization of the data auditing, licensing and traceability systems. In India, where the regulation of AI is still in its infancy, the ruling may be viewed as a good precedent, at least until the time courts start to grapple with the question of the legality of generative AI. This logic is obvious internationally; piracy remains a contaminant in the well and the challenge lies in ensuring that technological advancement does not come at a cost of creators’ rights.
Conclusion
The Anthropic ruling is a significant turning point in the judicial control of generative AI, and it serves as a reminder that fair use should not be construed as a blanket license for experimental use. It can be well be summarized in the case that the existence of the venues to collect training data has ceased to be disregarded as fair use and can be marked as a changing of direction in fair use jurisprudence, beyond a mere analysis of purpose and output, to a sense of provenance and legitimacy of the input itself.
The Bartz decision has far-reaching consequences, not only of law, but also of trend. Generative AI is at the heart of how we learn, interact and create, so the integrity of these systems should be determined not only by the quality of their products but also by the legality and morality of their training sequences. People must be able to trust in the product of AI as well as the way that it is constructed.
The lesson of this case is more than a lesson to the developer; it is a social lesson that ethical innovation should rest on a legal basis. The distinction between fair use and foul play is not only in how it is applied but also in acquisition. Hence, Law should be part and parcel of the design of AI—not an afterthought.
This blog is written by Anshika Gupta and Pulak Bisen, 3rd Year students, Rajiv Gandhi National University of Law, Punjab.