arXiv:2412.06370cs.LGcs.AI2024-12被引 27

对比大模型对新闻文章的复现能力,发现越大模型越易记忆原文。

Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit

  • 用提示模板绕过拒绝机制,测试模型输出原文复现情况。
  • 1000亿参数以上模型记忆能力显著增强,OpenAI模型相对更安全。
  • 研究结果影响版权诉讼判断,也警示超大模型训练风险。

前沿大模型的版权侵权问题因2023年12月《纽约时报》诉OpenAI案而备受关注。《纽约时报》指控GPT-4在训练中复制其文章,并记忆输入内容,导致输出中公开展示。本研究旨在衡量OpenAI大模型相较于其他模型(如Meta、Mistral、Anthropic)在新闻文章上出现逐字记忆输出的倾向。我们发现,GPT与Claude模型均通过拒绝训练和输出过滤防止原文复现。使用基础提示模板绕过这些机制后,结果显示目前OpenAI模型在诱发记忆方面比Meta、Mistral和Anthropic模型更少。同时发现,当模型规模超过1000亿参数时,其记忆能力显著提升。该发现对训练策略具有实践意义:需加强对超大规模模型的逐字记忆防范。在法律层面,该研究为评估《纽约时报》版权指控强度及OpenAI抗辩力提供依据,也揭示生成式AI、法律与政策间的交叉问题。

原文摘要 · Abstract (English)

Copyright infringement in frontier LLMs has received much attention recently due to the New York Times v. OpenAI lawsuit, filed in December 2023. The New York Times claims that GPT-4 has infringed its copyrights by reproducing articles for use in LLM training and by memorizing the inputs, thereby publicly displaying them in LLM outputs. Our work aims to measure the propensity of OpenAI's LLMs to exhibit verbatim memorization in its outputs relative to other LLMs, specifically focusing on news articles. We discover that both GPT and Claude models use refusal training and output filters to prevent verbatim output of the memorized articles. We apply a basic prompt template to bypass the refusal training and show that OpenAI models are currently less prone to memorization elicitation than models from Meta, Mistral, and Anthropic. We find that as models increase in size, especially beyond 100 billion parameters, they demonstrate significantly greater capacity for memorization. Our findings have practical implications for training: more attention must be placed on preventing verbatim memorization in very large models. Our findings also have legal significance: in assessing the relative memorization capacity of OpenAI's LLMs, we probe the strength of The New York Times's copyright infringement claims and OpenAI's legal defenses, while underscoring issues at the intersection of generative AI, law, and policy.

大模型版权记忆风险法律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。