提出统一攻击框架,可同时伪造和清除模型水印
Unified attacks to large language model watermarks: spoofing and scrubbing in unauthorized knowledge distillation
- 用对比解码提取水印文本,实现双向蒸馏攻击
- 在保持模型性能前提下,成功移除或伪造水印痕迹
- 揭示现有水印机制漏洞,适合安全与合规研究者
水印已成为应对大语言模型虚假信息和保护知识产权的关键技术。近期发现的‘水印放射性’表明,教师模型中的水印可通过知识蒸馏传递给学生模型。正向看,这可用于检测未经授权的知识蒸馏;但水印在未经许可的蒸馏下对擦除攻击的鲁棒性及防伪造能力仍不明确。现有攻击方法或需访问模型内部,或无法同时支持擦除与伪造。本文提出对比解码引导的知识蒸馏(CDG-KD)框架,可在未经授权蒸馏下实现双向攻击。该方法通过对比学生模型与弱水印参考输出,提取被污染或放大的水印文本,再通过双向蒸馏训练新学生模型,分别实现水印去除与伪造。大量实验表明,CDG-KD在保持模型通用性能的同时有效执行攻击。研究强调了开发具备鲁棒性与不可伪造性的水印方案的紧迫性。
原文摘要 · Abstract (English)
Watermarking has emerged as a critical technique for combating misinformation and protecting intellectual property in large language models (LLMs). A recent discovery, termed watermark radioactivity, reveals that watermarks embedded in teacher models can be inherited by student models through knowledge distillation. On the positive side, this inheritance allows for the detection of unauthorized knowledge distillation by identifying watermark traces in student models. However, the robustness of watermarks against scrubbing attacks and their unforgeability in the face of spoofing attacks under unauthorized knowledge distillation remain largely unexplored. Existing watermark attack methods either assume access to model internals or fail to simultaneously support both scrubbing and spoofing attacks. In this work, we propose Contrastive Decoding-Guided Knowledge Distillation (CDG-KD), a unified framework that enables bidirectional attacks under unauthorized knowledge distillation. Our approach employs contrastive decoding to extract corrupted or amplified watermark texts via comparing outputs from the student model and weakly watermarked references, followed by bidirectional distillation to train new student models capable of watermark removal and watermark forgery, respectively. Extensive experiments show that CDG-KD effectively performs attacks while preserving the general performance of the distilled model. Our findings underscore critical need for developing watermarking schemes that are robust and unforgeable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。