arXiv:2603.20957cs.CLcs.AI2026-03被引 6

微调让大模型泄露版权书内容,连法庭信誓旦旦的防护都失效。

Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

  • 用情节摘要微调模型,即可触发对版权书的原文复现。
  • 三款主流模型复现率达85%-90%,最长原文段落超460词。
  • 跨作者泛化且行业通用,暴露模型存储版权内容的根本漏洞。

前沿大模型公司反复向法院和监管机构保证,其模型不会存储训练数据副本。他们依赖强化学习人类反馈(RLHF)、系统提示和输出过滤等安全对齐策略,以阻止对受版权保护作品的原文复现,并以此作为应对版权侵权诉讼的法律辩护依据。本文揭示:微调可绕过这些防护机制——通过训练模型将情节摘要扩展为完整文本(一项适合商业写作助手的任务),我们使GPT-4o、Gemini-2.5-Pro和DeepSeek-V3.1在仅使用语义描述作为提示的情况下,复现高达85%-90%的未见版权书籍内容,单个原文片段超过460词。该提取能力跨作者泛化:仅在村上春树小说上微调,即可触发对超过30位无关作者的版权书原文复现。该现象不特指某作者或语料库:随机作者配对与公共领域数据微调也产生类似效果,而合成文本微调则几乎无提取,表明微调激活了预训练阶段的潜在记忆。三家来自不同厂商的模型在同一书籍区域复现率均≥0.90,指向全行业性漏洞。研究结果有力证明模型权重中存有受版权保护作品的副本,且微调后显现的安全失效,动摇了近期合理使用裁决所依赖的核心前提——即只要采取足够措施防止受保护表达复现,即可豁免责任。

原文摘要 · Abstract (English)

Frontier LLM companies have repeatedly assured courts and regulators that their models do not store copies of training data. They further rely on safety alignment strategies via RLHF, system prompts, and output filters to block verbatim regurgitation of copyrighted works, and have cited the efficacy of these measures in their legal defenses against copyright infringement claims. We show that finetuning bypasses these protections: by training models to expand plot summaries into full text, a task naturally suited for commercial writing assistants, we cause GPT-4o, Gemini-2.5-Pro, and DeepSeek-V3.1 to reproduce up to 85-90% of held-out copyrighted books, with single verbatim spans exceeding 460 words, using only semantic descriptions as prompts and no actual book text. This extraction generalizes across authors: finetuning exclusively on Haruki Murakami's novels unlocks verbatim recall of copyrighted books from over 30 unrelated authors. The effect is not specific to any training author or corpus: random author pairs and public-domain finetuning data produce comparable extraction, while finetuning on synthetic text yields near-zero extraction, indicating that finetuning on individual authors' works reactivates latent memorization from pretraining. Three models from different providers memorize the same books in the same regions ($r \ge 0.90$), pointing to an industry-wide vulnerability. Our findings offer compelling evidence that model weights store copies of copyrighted works and that the security failures that manifest after finetuning on individual authors' works undermine a key premise of recent fair use rulings, where courts have conditioned favorable outcomes on the adequacy of measures preventing reproduction of protected expression.

版权模型记忆微调漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。