将水印嵌入低维功能子空间,提升大模型版权保护的鲁棒性。
Functional Subspace Watermarking for Large Language Models
- 通过求解广义特征值问题提取稳定的功能子空间用于水印注入。
- 在多种攻击下检测准确率超现有最优方法,且保持模型性能不变。
- 适合需要强鲁棒性水印的大模型版权保护场景。
模型水印利用内部表示来保护大型语言模型(LLM)的所有权。然而,在微调、量化或知识蒸馏等实际模型修改过程中,这些特征不可避免地经历复杂失真,导致可靠提取极为困难。尽管已有大量关于模型端水印的研究,现有方法对参数级扰动仍缺乏足够鲁棒性。为此,我们提出功能子空间水印(Functional Subspace Watermarking, FSW),将所有权信号锚定在低维功能主干中。具体地,我们首先通过求解广义特征值问题提取稳定的函数子空间以进行水印注入,并引入自适应谱截断策略,在鲁棒性与模型效用间取得最佳平衡。此外,还加入向量一致性约束,确保水印注入不损害原始语义性能。在多种LLM架构和数据集上的大量实验表明,该方法在多种模型攻击下均实现更优的检测准确率与统计可验证性,其鲁棒性优于现有SOTA方法。
原文摘要 · Abstract (English)
Model watermarking utilizes internal representations to protect the ownership of large language models (LLMs). However, these features inevitably undergo complex distortions during realistic model modifications such as fine-tuning, quantization, or knowledge distillation, making reliable extraction extremely challenging. Despite extensive research on model-side watermarking, existing methods still lack sufficient robustness against parameter-level perturbations. To address this gap, we propose \texttt{\textbf{Functional Subspace Watermarking (FSW)}}, a framework that anchors ownership signals into a low-dimensional functional backbone. Specifically, we first solve a generalized eigenvalue problem to extract a stable functional subspace for watermark injection, while introducing an adaptive spectral truncation strategy to achieve an optimal balance between robustness and model utility. Furthermore, a vector consistency constraint is incorporated to ensure that watermark injection does not compromise the original semantic performance. Extensive experiments across various LLM architectures and datasets demonstrate that our method achieves superior detection accuracy and statistical verifiability under multiple model attacks, maintaining robustness that outperforms existing state-of-the-art (SOTA) methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。