Global Technology Editor

Generative AI has moved from a question of capability to a question of permission. The models can ingest vast libraries of text, images, and code; the harder issue is whether the law treats that ingestion as legitimate learning, unlawful copying, or something in between. Around the world, the answer is beginning to look less like a philosophical debate and more like an operational rulebook for the AI economy.

In the United States, the legal picture is being shaped by a mix of litigation and public guidance.[1][2] The U.S. Copyright Office released Part 2 of its multi-part report on Copyright and Artificial Intelligence on January 29, 2025, focusing on whether outputs created with generative AI can be copyrighted.[1][4] The office said Part 1 had addressed digital replicas and that a future Part 3 would cover training, fair use, and licensing.[1][4] That sequencing matters: the law is being assembled piece by piece, and companies are building products in the gaps between those pieces.

Authors Guild filed a class action in Manhattan federal court against OpenAI on behalf of prominent writers, arguing that training on their books was unlawful.[2][12] Other copyright owners have brought similar claims against large AI providers, while defendants have generally said internet-scraped training data falls within fair use.[2][6][10][12] A separate line of cases has made the risk harder to ignore.[6][10] In disputes over training on copyrighted material, courts have increasingly been asked whether a model’s transformation of text into learned parameters is enough to excuse copying, or whether the source market still matters.

If a model is trained on expressive works and then generates outputs that compete with those works, the legal theory weakens.[5][10] If the training is noncommercial research, or if the system does not expose protected expression in its outputs, the fair-use argument becomes stronger.[5][8] One legal analysis of the Copyright Office’s 2025 report captured the core tension: there is unlikely to be one answer for all uses, only a spectrum that runs from research to direct market substitution.[5] For AI builders, that spectrum is now a business constraint, not a footnote.

Japan has chosen a different architecture, and that contrast is revealing.[7][9] Under Article 30-4 of Japan’s Copyright Act, uses that do not aim to enjoy the thoughts or feelings expressed in a work can fall under a flexible rights-restriction rule.[3][7][9][11] Cultural Affairs materials describe this framework as designed to smooth the use of copyrighted works for innovation where market harm is limited.[7] In practice, that has made Japan an important reference point for companies seeking clearer rules for AI training. It also shows how national systems can diverge even when they are answering the same technical question: how much copying is necessary for machine learning?

The important detail is that neither the U.S. nor Japan is really debating learning in the abstract.[3][7][9][11] Both are debating the conditions under which copying for computational analysis should be tolerated.[3][7][9][11] In Japan, the legal discussion often turns on whether the use is for information analysis rather than enjoyment of the work itself.[3][9][11] In the United States, the discussion is still filtered through fair use and the four-factor test, which gives judges room to weigh purpose, nature, amount, and market effect.[5][6][10] For developers, that means the compliance burden is not simply about where the servers sit. It is about dataset provenance, licensing posture, output controls, and the commercial proximity of the model to the works it ingests.

This is where the policy question becomes a capital question. Large model builders can spread legal risk across enormous revenue bases, negotiate licenses, or absorb the cost of protracted litigation. Smaller companies cannot as easily do any of those things. If training data becomes a licensed input rather than an open resource, the cost structure of model development changes.[5][8] That would not stop AI, but it would shift advantage toward firms with money, legal teams, and distribution. In that sense, copyright is becoming part of the infrastructure stack, much like cloud compute or advanced chips.

There is still important uncertainty, and it should be stated plainly. The provided materials do not settle how every court will treat training, nor do they prove that all models trained on copyrighted works are unlawful or lawful.[1][5][6][10] They do suggest that the old habit of describing AI training as a wholly settled form of fair use is no longer credible.[5][6][10] What would change the reading? A major appellate decision, a final U.S. Copyright Office position on training, a market-wide licensing regime, or a legislative intervention in either the United States or another large market.[1][5][7] Any of those would redraw the practical boundary.

The international dimension is easy to underestimate. If the United States tightens around fair use while Japan preserves a more explicit training pathway, companies may route development, partnerships, and compliance structures accordingly.[3][7][9][11] That would make copyright law part of AI competition between jurisdictions, not just between plaintiffs and defendants. The real contest is not over whether machines can read. It is over who pays when they do, and which legal regime gets to define the price of imitation. That is the issue to watch, because it will shape the terms on which generative AI becomes a durable part of the global economy.