arXiv:2501.09284cs.AIcs.CR2025-01被引 5

为LoRA权重设计可抵抗多种攻击的嵌入式水印技术

SEAL: Entangled White-box Watermarks on Low-Rank Adaptation

  • 在可训练的LoRA权重间嵌入不可训练的密钥矩阵
  • 水印与权重联合训练,无性能损失且抗移除/混淆攻击
  • 适合需版权保护的模型微调成果共享场景

LoRA及其变体因高效简便,已成为训练和共享特定任务大模型的主流方法。然而,针对LoRA权重的版权保护,尤其是基于水印的技术仍研究不足。为此,我们提出SEAL(SEcure wAtermarking on LoRA weights),一种适用于LoRA的通用白盒水印方案。SEAL在可训练的LoRA权重间嵌入一个秘密、不可训练的矩阵作为所有权凭证,并通过训练将该凭证与权重纠缠,无需额外损失函数。训练完成后,密钥被隐藏并随微调权重分发。实验表明,SEAL在常识推理、文本/视觉指令微调及文生图任务中均无性能下降。同时,其对移除、混淆和模糊等已知攻击具有鲁棒性。

原文摘要 · Abstract (English)

Recently, LoRA and its variants have become the de facto strategy for training and sharing task-specific versions of large pretrained models, thanks to their efficiency and simplicity. However, the issue of copyright protection for LoRA weights, especially through watermark-based techniques, remains underexplored. To address this gap, we propose SEAL (SEcure wAtermarking on LoRA weights), the universal whitebox watermarking for LoRA. SEAL embeds a secret, non-trainable matrix between trainable LoRA weights, serving as a passport to claim ownership. SEAL then entangles the passport with the LoRA weights through training, without extra loss for entanglement, and distributes the finetuned weights after hiding the passport. When applying SEAL, we observed no performance degradation across commonsense reasoning, textual/visual instruction tuning, and text-to-image synthesis tasks. We demonstrate that SEAL is robust against a variety of known attacks: removal, obfuscation, and ambiguity attacks.

LoRA水印版权保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。