学会誌・資料 Journal

日本知財学会誌 第21巻第1号掲載

画像生成AI におけるデータセットと知的財産をめぐる課題
Dataset Structure in Image Generation AI and IP Issues

渡邉 恵理子, 北野 孝太, 脇田 英
Eriko Watanabe, Kota Kitano, Suguru Wakita

日本知財学会誌Vol.21 No.1 p.36-44 (2024-9-20)
Journal of Intellectual Property Association of Japan Vol.21 No.1 p.36-44 (2024-9-20)

<要旨>
2022年以降,生成AIが急速に普及し,様々なメディアの生成に利用されているが,学習データの権利処理や生成されたコンテンツの知的財産権についての懸念が高まっている.本論文では,Stable Diffusion を例に,使用されるデータセットの構造や,再学習・補助モデルの拡張技術,また画像生成AIサービスが提供するモデルと学習データの公開状況などを整理する.つづいて,画像生成AIサービス・コミュニティサイトの現状を示し,著作権侵害とみなせる画像生成モデルのアップロード例が多数存在する等の課題を述べる.画像生成AI を含む生成AIは知的財産の活用を推進する一方で,権利者が把握できない利用が発生している.画像生成AIの社会実装をより発展的に進めるには,開発者と利用者だけでなく,AI生成物の品質を支える学習データの大元である著作者も含めた社会的な合意形成が必要と考えられる.
<Abstract>
Since 2022, generative AI has gained popularity for creating various media, raising concerns about the rights management of training data and intellectual property rights of generated content. This paper uses Stable Diffusion as a case study to outline the structure of datasets, techniques for retraining and extending auxiliary models. 
The publication status of models and training data provided by image generation AI services is also presented. Furthermore, the current state of image generation AI services and community sites is examined, with particular attention paid to issues such as the prevalence of models that potentially infringe copyrights. While generative AI facilitates the use of intellectual property, it often leaves rights holders unaware of its usage. In order to facilitate the social implementation of image generation AI, it is necessary to establish a social consensus. This should involve not only developers and users but also the original authors of the training data that underpin the quality of AI-generated content.

<キーワード>
画像生成AI, データセット, 知的財産, 再学習, 著作権侵害
<Keywords>
Image generation AI, Dataset, Intellectual property, Finetuning, Copyright infringement