arXiv:2602.12270econ.THcs.AI2026-02

AI生成作品是否侵权?新标准看它能否离开原训练数据。

Creative Ownership in the Age of AI

  • 用训练数据依赖性定义侵权:若无原作就无法生成,则算侵权。
  • 轻尾创作下,AI可自由生成;重尾创作下,需持续监管。
  • 适合法律界、AI伦理研究者和版权政策制定者阅读。

版权法关注新作品与现有作品是否“实质性相似”,但生成式AI能高度模仿风格而不复制内容,这已成为当前诉讼的核心。我们指出现有侵权定义在此场景下不适用,并提出新标准:若生成式AI输出无法在缺乏某作品训练数据的情况下生成,则构成侵权。为实现这一标准,我们将生成系统建模为将现有作品语料映射到新作品输出的闭包算子。根据该标准,若输出不侵犯任何现有作品,则为允许生成。我们的研究揭示了允许生成的结构特性,并发现显著的渐近二分现象:当有机创作呈轻尾分布时,对单个作品的依赖最终消失,监管不再限制生成;而当创作呈重尾分布时,监管将持续具有约束力。

原文摘要 · Abstract (English)

Copyright law focuses on whether a new work is "substantially similar" to an existing one, but generative AI can closely imitate style without copying content, a capability now central to ongoing litigation. We argue that existing definitions of infringement are ill-suited to this setting and propose a new criterion: a generative AI output infringes on an existing work if it could not have been generated without that work in its training corpus. To operationalize this definition, we model generative systems as closure operators mapping a corpus of existing works to an output of new works. AI generated outputs are \emph{permissible} if they do not infringe on any existing work according to our criterion. Our results characterize structural properties of permissible generation and reveal a sharp asymptotic dichotomy: when the process of organic creations is light-tailed, dependence on individual works eventually vanishes, so that regulation imposes no limits on AI generation; with heavy-tailed creations, regulation can be persistently constraining.

版权AI生成法律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。