arXiv:2502.15858cs.CYcs.AI2025-02被引 10

生成式AI训练数据版权风险被高估,现有法律豁免不适用。

Generative AI Training and Copyright Law

  • 指出生成式AI训练与文本数据挖掘本质不同,无法适用TDM例外
  • 揭示训练数据记忆现象会引发独立版权侵权问题
  • 呼吁学术界共建兼顾各方利益的AI训练规范

生成式AI训练需海量数据,通常通过网络爬取获取。但大量数据受版权保护,其使用可能构成侵权。美国开发者依赖“合理使用”原则,欧洲则普遍认为“文本与数据挖掘”(TDM)例外适用。本文基于跨学科研究指出,生成式AI训练与TDM存在根本差异,现有法律豁免并不成立。同时,训练数据记忆现象会独立引发版权争议。文章进一步探讨了如何在公共与企业研究中建立合规实践,并提出ISMIR可在推动多方共赢的生成式AI伦理框架建设中发挥作用。

原文摘要 · Abstract (English)

Training generative AI models requires extensive amounts of data. A common practice is to collect such data through web scraping. Yet, much of what has been and is collected is copyright protected. Its use may be copyright infringement. In the USA, AI developers rely on "fair use" and in Europe, the prevailing view is that the exception for "Text and Data Mining" (TDM) applies. In a recent interdisciplinary tandem-study, we have argued in detail that this is actually not the case because generative AI training fundamentally differs from TDM. In this article, we share our main findings and the implications for both public and corporate research on generative models. We further discuss how the phenomenon of training data memorization leads to copyright issues independently from the "fair use" and TDM exceptions. Finally, we outline how the ISMIR could contribute to the ongoing discussion about fair practices with respect to generative AI that satisfy all stakeholders.

版权法AI训练数据合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。