无需微调即可为大模型添加抗攻击水印,保护知识产权
Invariant-based Robust Weights Watermark for Large Language Models
- 基于模型不变性构建线性约束生成稳定水印值
- 在三种模型上对多种攻击均保持水印有效性
- 适合需版权保护的边缘部署大模型场景
随着大语言模型在数十亿资源受限的边缘设备上部署,知识产权保护日益重要。为应对恶意用户窃取模型的风险,本文提出一种无需重训练或微调的鲁棒水印方案。该方案为每个用户生成唯一密钥,并通过求解由模型不变性构造的线性约束,得到稳定的水印值。同时,在多用户场景中引入噪声机制以隐藏水印位置,防范共谋攻击。实验在Llama3、Phi3、Gemma三个主流模型上进行,结果表明该方法在微调、剪枝、量化、置换、缩放、可逆矩阵变换及共谋攻击等多种攻击下均表现出强鲁棒性。
原文摘要 · Abstract (English)
Watermarking technology has gained significant attention due to the increasing importance of intellectual property (IP) rights, particularly with the growing deployment of large language models (LLMs) on billions resource-constrained edge devices. To counter the potential threats of IP theft by malicious users, this paper introduces a robust watermarking scheme without retraining or fine-tuning for transformer models. The scheme generates a unique key for each user and derives a stable watermark value by solving linear constraints constructed from model invariants. Moreover, this technology utilizes noise mechanism to hide watermark locations in multi-user scenarios against collusion attack. This paper evaluates the approach on three popular models (Llama3, Phi3, Gemma), and the experimental results confirm the strong robustness across a range of attack methods (fine-tuning, pruning, quantization, permutation, scaling, reversible matrix and collusion attacks).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。