区分记忆与抄袭,为大模型版权评估提供新框架
We Should Separate Memorization from Copyright
- 提出将技术上的记忆与法律上的抄袭区分开来
- 指出传统重建方法误判了版权风险
- 倡导基于输出的风险评估,契合法律标准
基础模型的广泛应用带来了新的版权风险。当前技术与法律界对数据使用问题存在分歧,常因解读不同导致不同结论。现有研究多依赖传统重建技术,而这些方法并非为版权分析设计,导致记忆与复制被混为一谈。本文主张:数据科学中普遍研究的‘记忆’不应等同于‘复制’,也不应作为版权侵权的代理指标。我们区分出真正反映侵权风险的技术信号,以及仅体现合法泛化或高频内容的信号。基于此,提出应在输出层面采用基于风险的评估机制,使技术判断与既有版权标准一致,为研究、审计和政策制定提供更严谨的依据。
原文摘要 · Abstract (English)
The widespread use of foundation models has introduced a new risk factor of copyright issue. This issue is leading to an active, lively and on-going debate amongst the data-science community as well as amongst legal scholars. Where claims and results across both sides are often interpreted in different ways and leading to different implications. Our position is that much of the technical literature relies on traditional reconstruction techniques that are not designed for copyright analysis. As a result, memorization and copying have been conflated across both technical and legal communities and in multiple contexts. We argue that memorization, as commonly studied in data science, should not be equated with copying and should not be used as a proxy for copyright infringement. We distinguish technical signals that meaningfully indicate infringement risk from those that instead reflect lawful generalization or high-frequency content. Based on this analysis, we advocate for an output-level, risk-based evaluation process that aligns technical assessments with established copyright standards and provides a more principled foundation for research, auditing, and policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。