The text and images generated by AI within the blink of an eye appear to be new creations. However, they are actually products of enormous amounts of collected training material. This raises concerns about the ethical use of data, consent, and compensation for the human-centred work behind it. The UK government highlighted this concept through its copyright report of 2026, which showcased how frontier AI models now require billions of inputs, including works that are copyright-protected. This data includes books, journals, photographs, articles, websites, illustrations, code, and other publicly accessible material.
Even though data is available online and freely accessible, this does not mean that the creator has surrendered all their rights. This creates a tension between access and ownership. When developing AI, datasets are treated as infrastructure for building technology. For companies, this becomes one small point within a much larger scale of development. However, for creators, these same datasets represent years of creative labour and intellectual effort.
This poses an important question: if human creativity provides the raw material, should AI have the power to use that material without discussing compensation, consent, and copyright?
Copyright was built around the idea of copying, but AI has complicated the equation.
Traditionally, copyright was put in place to give creators autonomy and control over how their craft or ideas are reproduced and used. However, with the advent of vastly capable AI models, this picture has become more complicated. AI systems are trained on datasets, not necessarily by reproducing an entire book or article, but by analysing patterns, similarities, and relationships within enormous amounts of material without necessarily asking for individual consent.
AI models are trained by analysing enormous quantities of material and learning from statistical patterns. The UK government has described AI systems as translating training information into statistical representations of the world. Thus, many wonder whether training AI models on already copyright-protected material is equivalent to copying that material, or whether it is closer to learning from information that already exists and is open to the public.
Different institutions answer this question differently. The U.S. Copyright Office, for instance, received more than 10,000 public comments during its AI copyright inquiry. An AI training report was also issued in 2025, examining fair use, potential liability, and licensing. This shows that there is no one universal answer, but the debate over AI, creative liberty, and autonomy is still continuing.
For creators, writers, illustrators, photographers, and many others — several concerns have emerged, with consent being among the top priorities.
Was permission ever given for AI systems to borrow from and learn from their work? Should proper credit be provided whenever their work is used? And, lastly, compensation: should creators receive payment or another form of incentive when their work and labour contribute to AI development?
The Australian Society of Authors conducted a survey in 2025. The statistics showed that 98% of respondents believed AI companies should seek permission before using illustrations or texts for training. Alongside this, 92% wanted compensation for work that had already been used without consent or without informing them, while 93% expected compensation where consent was given through a proper licensing agreement.
For a writer, an article represents hours of research, editing, and careful formation of sentences. For an artist, an illustration may represent years of developing a particular technique. For a journalist, a report may be the result of several interviews, data collection, and fieldwork.
When we add thousands of such experiences together, the picture becomes clearer. Creators deserve to know where the line between inspiration and extraction lies.
India's ANI versus OpenAI case provides an important example. The Delhi High Court in July 2026 made a substantive finding concerning OpenAI's use of ANI's news content to train ChatGPT, holding that the use fell within a fair-dealing exception for research.
This becomes particularly important for journalists because news organisations produce enormous quantities of copyright-protected material.
India's copyright framework already contains fair-dealing exceptions under Section 52. But the larger policy question still remains: should AI training be treated like research, or should creators be able to opt out of the process? Should AI companies instead be required to license the material they use? Could collective licensing systems help address these future systemic issues?
The debate is gradually moving from AI versus creators towards finding possible mechanisms for coexistence.
The UK government's March 2026 report showcased how licensing, transparency, technical standards, ethical enforcement, and mechanisms allowing rights holders greater control over their work could help creators and AI developers find common ground. The report identified different approaches, including mandatory licensing, data-mining exceptions, and exceptions combined with opt-out and transparency mechanisms.
AI licensing is already becoming a market. The Publishers Association in 2026 announced that publishers are actively licensing content for training and building AI models. This suggests that human-created, high-quality material itself is becoming a more valuable asset for AI.
Now, the dynamic has shifted. Instead of simply asking how AI should use human work ethically, or whether it should use it at all, the emerging question is how that use should be negotiated and paid for.
There is no single universal owner of AI or the data that it utilises. Some material might be copyrighted, some might be obtained from the public domain or acquired through open-source licences. Other information might have been collected with explicit permission from original owners.
Therefore, instead of asking who owns the data, a better question would be: who determines the terms under which data can be accessed and used?
AI developers need clearer ethical guidelines to use datasets responsibly without generalising the rights attached to them. At the same time, creators need copyright protection and economic models that ensure their work can continue to have value rather than simply becoming part of a pool of material used by AI systems.
The future of AI does not solely depend on how much data machines can consume. It also depends on whether human-centred information can continue to thrive alongside AI expansion.
References: