提出可渐进删除版权内容的LLM去学习框架,解决模型持续侵权问题。
Investigating the Feasibility of Mitigating Potential Copyright Infringement via Large Language Model Unlearning
- 用任务向量定位并移除版权内容对应的权重更新
- 在多阶段去学习中平衡删版权与保留通用语言能力
- 适合关注AI版权合规与模型可解释性的研究者
预训练大语言模型虽表现出色,却可能学习并生成受版权保护的内容,引发重大法律与伦理问题。在真实场景中,模型所有者需随时间陆续处理内容删除请求。现有研究未系统探索此类连续去学习机制。为此,我们提出稳定渐进去学习(SSU)框架,通过任务向量识别并移除对应版权内容的权重更新,结合随机标签损失提升去学习效果,并利用基于梯度的权重显著性调整目标参数以保留通用知识。大量实验表明,SSU在部分情况下实现了去学习效果与通用语言能力间的有效权衡,优于现有基线方法,但并非万能解法。
原文摘要 · Abstract (English)
Pre-trained Large Language Models (LLMs) have demonstrated remarkable capabilities but also pose risks by learning and generating copyrighted material, leading to significant legal and ethical concerns. In a potential real-world scenario, model owners may need to continuously address copyright infringement in order to address requests for content removal that emerge at different time points. One potential way of addressing this is via sequential unlearning, where copyrighted content is removed sequentially as new requests arise. Despite its practical relevance, sequential unlearning in the context of copyright infringement has not been rigorously explored in existing literature. To address this gap, we propose Stable Sequential Unlearning (SSU), a novel framework designed to unlearn copyrighted content from LLMs over multiple time steps. Our approach works by identifying and removing specific weight updates in the model's parameters that correspond to copyrighted content using task vectors. We improve unlearning efficacy by introducing random labeling loss and ensuring the model retains its general-purpose knowledge by adjusting targeted parameters with gradient-based weight saliency. Extensive experimental results show that SSU sometimes achieves an effective trade-off between unlearning efficacy and general-purpose language abilities, outperforming existing baselines, but it's not a cure-all for unlearning copyrighted material.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。