arXiv:2501.15478cs.CRcs.LG2025-01被引 3

为LoRA设计黑盒水印,有效追踪其非法使用。

LoRAGuard: An Effective Black-box Watermarking Approach for LoRAs

  • 提出阴阳水印机制,支持加法与反向操作下的水印验证。
  • 在多LoRA融合场景下,水印识别成功率接近100%。
  • 适用于语言与扩散模型,适合需要版权保护的研究者。

LoRA(低秩适应)在大模型参数高效微调中取得显著成功。训练后的LoRA矩阵可通过加法或反向运算与基础模型结合,提升下游任务性能。然而,未经许可使用LoRA生成有害内容的问题凸显了追踪其使用需求。自然的解决方案是将水印嵌入LoRA以检测非法滥用。但现有方法在多个LoRA组合或应用反向操作时表现不佳,因这些操作会严重削弱水印效果。本文提出LoRAGuard,一种新型黑盒水印技术,用于检测LoRA的非法使用。为支持加法与反向操作,我们提出阴阳水印机制:反向操作时验证阴水印,加法操作时验证阳水印。此外,我们设计基于影子模型的水印训练方法,在多LoRA集成场景中显著提升有效性。在语言与扩散模型上的大量实验表明,LoRAGuard实现近100%的水印验证成功率,表现出强鲁棒性。

原文摘要 · Abstract (English)

LoRA (Low-Rank Adaptation) has achieved remarkable success in the parameter-efficient fine-tuning of large models. The trained LoRA matrix can be integrated with the base model through addition or negation operation to improve performance on downstream tasks. However, the unauthorized use of LoRAs to generate harmful content highlights the need for effective mechanisms to trace their usage. A natural solution is to embed watermarks into LoRAs to detect unauthorized misuse. However, existing methods struggle when multiple LoRAs are combined or negation operation is applied, as these can significantly degrade watermark performance. In this paper, we introduce LoRAGuard, a novel black-box watermarking technique for detecting unauthorized misuse of LoRAs. To support both addition and negation operations, we propose the Yin-Yang watermark technique, where the Yin watermark is verified during negation operation and the Yang watermark during addition operation. Additionally, we propose a shadow-model-based watermark training approach that significantly improves effectiveness in scenarios involving multiple integrated LoRAs. Extensive experiments on both language and diffusion models show that LoRAGuard achieves nearly 100% watermark verification success and demonstrates strong effectiveness.

LoRA水印版权保护黑盒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。