Japan’s Cabinet Office has proposed a revised Principle-Code for generative artificial intelligence (AI) businesses that sets out principles on intellectual property (IP) protection, including avoiding crawling so-called pirate sites, respecting access restrictions such as paywalls, and increasing transparency around how businesses manage IP risks. The draft would use a ‘comply or explain’ approach, under which Generative AI businesses would either follow the principles or explain publicly why they do not.
The proposal could give rights holders more information about the models, training data and data-collection methods used by generative AI businesses, while leaving Japan’s existing copyright framework largely intact. It also puts transparency around data collection into a non-binding governance framework that would cover Japanese businesses and foreign businesses whose Generative AI systems or services are available in Japan.
The revised “Principle-Code for Protection of Intellectual Property and Transparency for the Appropriate Use of Generative AI” was discussed by the Study Group on Intellectual Property Rights in the AI Era on August 18. The Cabinet Office’s meeting materials identify it as a revised draft, with changes made after a public consultation that ran from December 26, 2025 to January 26, 2026.
The proposed code would apply to generative AI developers and providers, including businesses outside Japan whose generative AI systems or services are provided in Japan or made available to Japanese nationals. It is intended to balance the development of generative AI with the protection of intellectual property rights and greater transparency for rights holders and users.
Protecting copyrighted works during AI development
The draft says generative AI businesses should establish principles for protecting IP rights and clarify responsibility for implementing them. It also proposes that businesses review these principles at least once a year and publish their substance.
The draft says businesses should establish processes to ensure that their use of data to develop and train generative AI does not infringe others’ intellectual property rights. It calls on them to respect access restrictions, including paywalls, and to use crawlers that follow machine-readable instructions such as robots.txt. It also asks them to endeavour to avoid crawling so-called pirate sites, to disclose their crawler measures for each user agent, and to give notice when those measures change.
That puts the acquisition and use of training data within the proposed governance framework. The principles therefore address how businesses obtain training material, not only what their models generate.
Training and output safeguards
The draft also proposes that businesses retain training-related logs for a certain period. Where possible, it asks them to take technical measures to prevent infringing outputs. As far as possible, it asks them to use measures such as digital watermarking and C2PA to verify content origin and provenance.
Businesses would also be expected to establish contact points for rights holders, make the requirements for an approach as clear as possible, and keep records of their responses. They would also be expected to tell users of their AI not to use outputs that appear to infringe.
More transparency around training data
The proposed framework also sets out ways for rights holders to seek information about the use of their works in AI development.
The draft does not require businesses to release every individual item of training data publicly. It does, however, contemplate public disclosure of specified information about models, training data and collection methods, alongside a mechanism for a rights holder pursuing a legal remedy to ask whether a specific URL or identifier they name was used in training or validation, limited to what the business can readily access and confirm. AI users would have an equivalent mechanism in relation to their own outputs. It also recognises limits where the information is proprietary, including trade secrets.
Japan has also examined this issue through its broader IP policy. Its intellectual property strategy materials identify transparency around training data and the relationship between AI development and copyrighted works as areas requiring further attention.
The proposal comes as the legal treatment of AI training remains under active discussion internationally. MediaNama has previously reported on the copyright questions surrounding generative AI training, including when using copyrighted material to train a model can amount to infringement.
Japan is particularly relevant to that debate because Article 30-4 of its Copyright Act permits certain uses of copyrighted works for purposes such as data analysis, including AI training, subject to specified conditions. A March 2024 document adopted by a subcommittee of the Council for Cultural Affairs said Article 30-4 can permit the use of copyrighted works for AI development and other data-analysis purposes without permission from the copyright holder where the statutory conditions are met. That document also said the exception does not apply where there is a purpose of enjoying the work, that it can fail where material is taken by circumventing access restrictions or from a paid database, and that businesses should strictly refrain from deliberately collecting from known pirate sites – the same ground the draft code now covers.
Japan is proposing governance around that copyright framework
The draft code does not amend Japan’s existing copyright rules. Instead, it proposes governance and disclosure practices for training data, IP protection and related safeguards.
That distinction is visible in the proposed “comply or explain” model. The framework is presented as a non-binding code rather than a new statutory obligation backed by a penalty.
Japan also has a broader AI governance framework under the Act on Promotion of Research and Development, and Utilization of Artificial Intelligence-related Technology. The law was promulgated on June 4, 2025, and came fully into force on September 1, 2025. The government adopted an AI Basic Plan under that framework by Cabinet decision on December 23, 2025. The IP principle code sits within that wider AI policy effort rather than amending the Copyright Act.
The code is still a draft
The August 18 document is not a law or regulation. Examples include respecting access restrictions, avoiding pirate sites, retaining training-related logs, creating channels for rights holders and using measures such as watermarking and C2PA.
The next question is how the government expects these principles to work in practice. That will include balancing rights-holder requests with trade secrets, security concerns and the technical difficulty of tracing individual works through large training datasets.
Also read: