THE LEGALITY OF TRAINING AI MODELS ON COPYRIGHTED MATERIAL IN INDIA: A CLOSER LOOK AT ANI V. OPENAI

On 24 July 2026, the Delhi High Court delivered the first substantive Indian ruling on whether an AI company can copy copyrighted material to train a large language model without a licence. Justice Amit Bansal refused ANI Media’s application for an interim injunction against OpenAI, holding that storing ANI’s news reports to train ChatGPT prima facie falls within India’s fair dealing exception. The order decides an interlocutory application, not the suit, and its findings are expressly provisional. Even so, arriving four days after a California court approved Anthropic’s USD 1.5 billion settlement with book authors, it places India inside a global argument about who pays for the material AI models learn from.

How the Dispute Reached the Court

  • ANI, one of India’s largest news agencies, sued OpenAI in November 2024 on two distinct grounds: (i) that OpenAI scraped and stored its published reports to train ChatGPT without permission or payment, and (ii) that ChatGPT fabricated stories and attributed them to ANI, damaging its reputation. ANI had offered to license its archive, but OpenAI declined.
  • OpenAI argued that facts and features like “news of the day” cannot be monopolised, that training a model is closer to reading than to copying, that publishers can block crawlers if they choose, and that the training copies were made on servers outside India and so beyond the reach of Indian copyright law. The Court appointed two amici curiae and allowed interventions by the Digital News Publishers Association and the Federation of Indian Publishers. The 135-page judgment followed some 32 hearings.

The Four Questions and How the Court Answered them

  • Can an Indian court decide this? The Court found a sufficient connection with India. ANI is based in Delhi, OpenAI offers its services to users in India, and the allegedly infringing outputs were generated here. It rejected OpenAI’s argument that the location of its servers placed the dispute beyond Indian copyright law, reasoning that otherwise infringement could be insulated from domestic law simply by locating servers abroad. The Court also treated copies made on foreign servers as part of a continuous chain of acts originating in India and therefore capable of attracting Indian copyright law.
  • Is storing training data “reproduction”? OpenAI argued that its models learn statistical patterns from training data and that the resulting model weights do not contain copies of the underlying works. The Court took a more straightforward view: electronically storing a literary work, whether temporarily or permanently, amounts to reproduction under Section 14 of the Copyright Act, 1957. The purpose for which the copy is made does not change that characterisation; the separate question is whether the copying is protected by an exception.
  • Does fair dealing excuse it? Prima facie yes, as “private or personal use, including research” under Section 52(1)(a) of the Copyright Act, 1957. The Court held that a commercial purpose does not, by itself, defeat the exception. Unlike other parts of Section 52, this provision contains no express non-commercial limitation, and the training corpus remained internal rather than being made available to users. It also interpreted “research” as capable of extending to machine learning. The more debatable point is whether training a commercial product can properly be characterised as “private” use simply because the underlying corpus remains internal.
  • Do the outputs infringe? Not on this record. ANI’s examples post-dated OpenAI’s stated training cut-offs, so they did not establish memorisation. The Court instead attributed the responses to live retrieval and found no substantial similarity with ANI’s articles as whole works. It also noted that ANI’s prompts were designed to elicit protected content. The false attribution claim remains to be decided at trial. 

How this Compares with the Leading U.S. Decisions

  • Bartz v. Anthropic. In June 2025 Judge Alsup held that training on lawfully acquired books was fair use and transformative, but that downloading and keeping some seven million pirated books was not. The case settled for USD 1.5 billion, roughly USD 3,000 per work across about 500,000 works, with final approval granted on 20 July 2026. The critical distinction was the provenance of the training material, rather than the act of training itself.
  • Kadrey v. Meta. The day after the Bartz judgment, Judge Chhabria likewise found Meta’s use of copyrighted works for AI training to be transformative. But the authors lost largely because they failed to establish meaningful market harm, including that Meta’s models would dilute the market for their works. The ruling therefore turned heavily on the evidentiary record, rather than establishing that AI training is generally protected by fair use.
  • Thomson Reuters v. Ross Intelligence. Fair use failed where a non-generative research tool was trained on Westlaw headnotes to compete with Westlaw itself, now on appeal before the Third Circuit. Separately, the consolidated OpenAI proceedings in New York survived dismissal on the output claims in October 2025, and discovery continues.
  • US fair use allows courts to weigh four factors case by case, with transformative use and market harm central to the decisions. Indian fair dealing, by contrast, protects only specified purposes, requiring the Delhi High Court to fit AI training within “private or personal use, including research.” The result may resemble Bartz and Kadrey, but the legal route is narrower. In India, whether AI training remains protected will depend on how far courts are prepared to stretch those statutory categories to accommodate commercial machine learning.

What it means in Practice

  • For AI developers, it means breathing room, but not a safe harbour. This is a single-judge interim order, appealable, on prima facie findings. Nothing in it protects pirated corpora, paywall circumvention, copying in breach of licence, retrieval systems that reproduce protected text, or verbatim outputs. Documented provenance, deletion policies, memorisation testing and output filters remain the practical defence.
  • For publishers and rights owners, the evidentiary burden is now clearer. Showing that copyrighted material was used for training may not, by itself, establish infringement. Claimants will need to connect specific works to the relevant model and demonstrate infringing reproduction or outputs, supported by evidence of resulting market harm. The ANI ruling also suggests that carefully documented licensing practices and rights reservations may strengthen a rights owner’s position, particularly where compensation for use of content is sought. 

Conclusion

The ruling is a landmark, but it decides less than the headlines suggest. It settles, provisionally, that ingesting publicly accessible content into a training pipeline is not by itself infringement in India, while leaving open what a model produces, where its data came from, and whether creators are paid. For clients on either side, the operative question has shifted from “was my content used” to “can I prove what came out, and what it cost me” which is an evidence and documentation problem before it is a legal one.

Authors: Shantanu Mukherjee, Varun Alase

Leave Us A Message

Cookie Consent with Real Cookie Banner