Abstract
Generative Artificial Intelligence (GenAI) models can create high-resolution artworks, complex pieces of software code and intricate musical compositions in mere seconds. But it is based on teaching machines colossal databases of human-generated creations, and this is a major conflict between tech and IP. This paper explores the relationship between AI and copyright law within the Indian legal framework, and compares the domestic copyright regulations with international treaty obligations. This study brings to light the legal gray area concerning the issue of AI authorship and the lack of consent in data ingestion, by examining Section 2(d)(vi) of the Indian Copyright Act, 1957, and certain judicial principles like Eastern Book Company v. D.B. Modak. The paper builds on contributions from the legal series Guilty Minds (“Aalaap”) and landmark international litigation such as Getty Images Ltd v Stability AI, and highlights the consequences of algorithmic micro-sampling for human artists. Moreover, it assesses the national requirements against international requirements such as the Berne Convention’s Three Step Test and the TRIPS Agreement. The paper argues for a balanced approach to the regulation of scraping, one that neither overrules statutory exceptions for T&DM nor fails to regulate them at all, and that the classification of copyright protection is to be fine-tuned to provide adequate protection for databases while also addressing the concerns of both copyright holders and scrapers.
1. Introduction: The Duality of Artificial Intelligence in Creative Ecosystems
AI tools, specifically Large Language Models and diffusion-based generative platforms, have evolved from administrative tools to self-generating literary, artistic, and musical content creators. These tools produce results similar to human skill, judgment and aesthetic taste. While this technology is a great driver for efficiency and innovation in sectors, it is important to note that the operation of this technology, inclusive of ingestion, analysis and synthesis of protected human creations, is a powerful catalyst.
This is one of the main dilemmas that faces Intellectual Property Rights (IPR): to stimulate technological progress while respecting the economic and moral rights of human creators. The removal of regulatory barriers and fair compensation for commercial extraction of human-authored works is detrimental to artists, authors and musicians. On the other hand, blanket bans or legally making data access difficult, on the one hand stifles technological development and on the other deprives emerging markets of technical growth.
There will need to be clear, enforceable rules in a regulatory framework that balance these competing interests. This paper provides a framework for legal protection of creators and technological development from the statutory provisions of the Indian Copyright Act, 1957, judicial precedents, real-life illustrations and international treaties.
2. Authorship and Ownership Under the Indian Copyright Act, 1957
2.1 Statutory Interpretation of Section 2(d)(vi)
The big challenge with providing copyright protection to AI-created works is that the copyright system was built with a human-centric perspective. According to the Indian Copyright Act, 1957, which is a statutory provision, Section 13(1), copyright is only for “original literary, dramatic, musical, and artistic works.”1 The definition of the ‘author’ under Section 2(d) varies according to the creative sectors that it addresses.2
To cover computer-generated works, Parliament added Section 2(d)(vi) to the Copyright Act in 1994, which makes the computer user the author, as defined by the Copyright Act. Traditional software tools, such as graphic design software, could be applied simply as a means of doing something; but with modern Generative AI, this does not work. Once a user types in a natural language question, the neural network chooses the distribution of the pixels or the colour palette or the sentence structure on its own.
This leaves uncertainty in law as to who actually “causes” the work to be created:
- The Prompter: The person who is giving instructions. But, under existing copyright doctrine, there is no copyright protection for basic ideas or instructions because of the idea-expression dichotomy.3
- AI Developer: The programmers that created the model. However, the output of end-users cannot be anticipated or controlled by the developers.
- The AI System: AI systems do not have a legal personality under Indian law, nor do they have any rights to property.
2.2 Jurisprudential Standards of Originality
In addition, a work must meet the threshold of originality for it to be protected by copyright. In Eastern Book Company v. D.B. Modak (2008), the Indian jurisprudence on originality changed drastically.4 The apex court of India had rejected the English doctrine of “sweat of the brow” (which did not require a minimum standard of skill or judgment) in favour of the Canadian requirement of a “minimum degree of skill and judgment.” In order for a work to be original it must show personal choice, selection, and intellectual effort, not so trivial that it is merely a mechanical exercise.
Applying the EBC standard to AI content shows that AI-generated content is a “pure machine output” and lacks human intellectual choice. Algorithmic processes work on the basis of statistical probability and pattern matching in multi-dimensional vector spaces, not on the basis of conscious judgment. As a result, the output of fully autonomous AI falls short of the requirements of the statutory test for originality and does not qualify for copyright protection.
2.3 Administrative Precedents and Comparative Approaches
The statutory problem was evident in 2020 itself when the Indian Copyright Office, for a fleeting moment, registered an art piece titled ‘Canvas’, which credits an AI application called RAGHAV (Robust AI-Generative Artwork Apparatus) as a co-writer with the human creator. The Copyright Office then sent a withdrawal notice, pointing out that it had made a procedural mistake by assigning a copyright to a non-human subject.
This attitude is consistent with other court rulings worldwide:
- In Thaler v Perlmutter (2023), a US Federal Court clarified that for copyright protection to be afforded to an AI-generated work, it must be produced with the human author’s input.5
- US Copyright Office Guidelines (2023): affirmed that purely AI-generated works are in the public domain. If an AI-generated element is significantly edited, arranged or modified by a human, then the human part may be protected.
3. Inbound Infringement: Data Scraping, Algorithmic Sampling, and Legal Vacuums
In contrast to output ownership, where the focus is on who owns the output created by an AI model, inbound infringement focuses on AI model training and its legality. Generative models have to be built on massive amounts of data, trillions of tokens from books, scholarly articles, software repositories and artwork – a lot of which is not used with the explicit permission or financial compensation of the original rights holder.
3.1 The Limits of Section 52 Fair Dealing
An unlicensed or unauthorized copying/duplication of a protected work triggers a violation of copyright under Section 51 of the Indian Copyright Act. Many AI creators argue that it is “Fair Dealing” under sub-paragraph (a) of Section 52(1)6 that web data is being scraped.
The Indian legislation, however, is more restrictive than the open-ended “Fair Use” doctrine under the United States. Section 52(1)(a)7 of the Indian Act lists all of the fair dealing exceptions:
- Private/confidential use (including research);
- Criticism or review;
- Reporting of current events and public lectures.
Private research exemptions are not applicable to commercial AI entities that go by subscription or enterprise software. Their massive web scraping goes beyond statutory research limits. Further, as in the EU, under Articles 3 and 4 of the EU Digital Single Market Directive, as well as Section 29A of the UK Copyright, Designs and Patents Act 1988,8 India has no explicit statutory exception for Text and Data Mining (TDM).9
3.2 Global Precedents and Practical Case Illustrations
This is observable in the legal battle as well as in public discourse around algorithmic sampling, data ingestion, and copyright infringement. Globally, the development of large-scale copyright infringement cases, such as Getty Images v. Stability AI,10 has shown how high-tech companies can utilize billions of copyrighted image assets without permission to create synthetic images, which frequently contain modified sections and altered watermarks from the original content.
This is a structural conflict, which is best portrayed in the Indian legal drama Guilty Minds, Season 1, Episode 5 (“Aalaap”).11 The story here is about classical musicians suing a software developer for his new AI product that takes thousands of previously recorded works, samples them and stitches them back together in a second to create new compositions.
These examples illustrate two core mechanisms used in copyright’s impact with AI:
- Micro-Sampling and Technical Invisibility: Protected human expressions are broken down into sub-second expressions, acoustic vectors, or mathematical tokens via generative models. The developers say that no one track or image is copied exactly, so the resulting work is a uniquely created work that can’t be infringed upon or copied. The secondary layer is still based on the illegal copying and manipulation of human labour, as outlined in Section 51.
- Economic Displacement: AI software creates instant, automated versions of human work – versions that are based on the human artists’ mastery, but without having to pay a fee for licensing their work to the content databases.
4. International Treaty Obligations and Compliance
The internal IP policy of India should be consistent with international treaties under the auspices of the World Intellectual Property Organization (WIPO) and the World Trade Organization (WTO).
4.1 Applicable Treaty Standards
Article 2(1) of the Berne Convention for the Protection of Literary and Artistic Works (1886) protects literary and artistic works, and Article 9(2) provides a limitation on exceptions to national copyright.12
The Agreement on Trade-Related Aspects of Intellectual Property Rights (TRIPS, 1994) has incorporated Berne obligations under Article 9(1), and limits exceptions under Article 13 to avoid uncompensated commercial exploitation which would interfere with the normal processes of the economic marketplace.13
Furthermore, computer programs and database compilations are also considered as intellectual creations under Articles 4 and 5 of the WIPO Copyright Treaty (1996),14 and the underlying rights of content creators are not affected.
4.2 The International Three-Step Test Analysis
As per Article 9(2) of the Berne Convention and Article 13 of TRIPS, there are three cumulative conditions in which the member states would be able to introduce copyright exceptions:
- There are a few special cases where the exception does not apply;
- It does not interfere with a normal use of the work;
- It does not unduly burden the rights to be attributed to the author.
The blanket exemption for commercial AI data scraping is not logical in light of the second and third conditions. Mass scraping robs creators of licensing markets, a normal way of commercialising their content, and threatens the economic value of human authorship. Thus, any national statutory mechanism should include compensatory structures, otherwise it would not be compliant with the treaty.
5. Proposing a Balanced Regulatory Framework for India
India needs a structured regulatory regime to accelerate the adoption of technology and safeguard human creators in the socio-economic and legal environment.
5.1 Statutory Licensing and Collective Management for TDM
India should update the Copyright Act, 1957, by specifically adding a Statutory License for Text and Data Mining, instead of making blanket exceptions for all data scraping or assuming that there is no infringement.
This system allows AI developers to use works that are publicly available and copyrightable, without requiring any individual negotiation, but instead paying a prescribed “statutory royalty.” The administration of royalty collection and disbursement should be done by the Collective Management Organisations (CMOs) like the Indian Reprographic Rights Organisation (IRRO) and the Indian Performing Right Society (IPRS).
5.2 Mandatory Transparency, Auditability, and Legislative Reform
AI builders must be legally obligated to have verifiable operational standards under the proposed Digital India Act,15 and amendments to the Information Technology Rules:
- Transparency: To remove transparency opacity, developers should record summary logs and provenance information of copyrighted material used in the development of the model.
- De Minimis Exceptions: Minor, trivial copying (De Minimis Non Curat Lex): Traditional copyright law allows for minor, trivial copying to be excluded. But when it’s automated system-wide, with high volumes of micro-samples, it has a big impact. Legislation needs to make clear that training in the commercial market through systematic micro-extraction is not de minimis and therefore is actionable under Section 51.
- Watermarking and Provenance: Require cryptographic metadata and visual/auditory watermarking of outputs created using AI to track and verify content authenticity and prevent market dilution.
5.3 Tiered Approach to Copyrightability
The Indian Copyright Office should provide for a three-part classification of computer-generated works to overcome the ambiguities in Section 2(d)(vi):
- Pure AI Outputs (Unassisted): Texts created solely by AI models using natural language prompts without human creativity will be in the public domain.
- Human-Assisted AI Works: For works that contain some elements of arrangement, selection, editing, or structural suggestion, but incorporate AI-generated elements such as prompts, where the EBC v. Modak standard of “minimum degree of skill and judgment” is met, the copyright is limited to the human element of such works.
- Category C — AI as an Auxiliary Tool: All traditional copyrighted creations by humans are fully protected by traditional copyright laws.
6. Conclusion
Artificial Intelligence is a step further into the future, but it is a step in the wrong direction for humanity. India’s Copyright Act, 1957 needs some statutory amendments to regulate Generative AI. This can be done by explicitly defining the term in Section 2(d)(vi) of the Copyright Act, enacting a statutory licensing scheme for the commercial use of Text and Data Mining, adopting transparency measures for datasets as per the Digital India Act to deter micro-sampling, and conforming to international treaties on copyright such as the Berne Three-Step Test, in order to create a legal ecosystem that promotes and protects India’s domestic innovation while balancing the economic and moral rights of content creators.
Footnotes
- Copyright Act 1957, s 13(1). ↩
- Copyright Act 1957, s 2(d)(vi). ↩
- R.G. Anand v Delux Films AIR 1978 SC 1613. ↩
- Eastern Book Company v D.B. Modak (2008) 1 SCC 1. ↩
- Thaler v Perlmutter (DDC 2023) Civil Action No. 22-1564 (BAH). ↩
- Copyright Act 1957, s 51. ↩
- Copyright Act 1957, s 52(1)(a). ↩
- Directive (EU) 2019/790 of the European Parliament and of the Council on copyright and related rights in the Digital Single Market [2019] OJ L 130/92, art 3, art 4. ↩
- Copyright, Designs and Patents Act 1988 (UK), s 29A. ↩
- Getty Images (US) Inc v Stability AI Inc (D Del 2023) Civil Action No. 23-135. ↩
- Guilty Minds (Season 1, Episode 5: “Aalaap”, Amazon Prime Video 2022). ↩
- Berne Convention for the Protection of Literary and Artistic Works (1886) art 2(1), art 9(2). ↩
- Agreement on Trade-Related Aspects of Intellectual Property Rights (1994) art 9(1), art 13. ↩
- WIPO Copyright Treaty (1996) art 4, art 5. ↩
- Ministry of Electronics and Information Technology (MeitY), Government of India, Draft Digital India Act Concept Note (2023). ↩