arXiv:2512.04044cs.LGcs.AI2025-12

提升开放权重大模型水印的质量与可检测性平衡

MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking

  • 以高斯水印为奖励信号,通过策略微调实现细粒度权重更新
  • 在不损失生成质量前提下,检测率接近推理时水印效果
  • 对改写和微调攻击鲁棒,跨数据集泛化能力强

水印旨在生成文本中嵌入隐藏信号,仅凭密钥可可靠检测。开放权重语言模型使此类水印面临严峻挑战,因现有依赖推理时干预的方法在权重公开后无法强制执行。现有开放权重水印技术如GaussMark,通常通过微调模型权重实现,但要达到与推理时水印相当的检测能力,需显著扰动权重,导致生成质量下降。本文提出MarkTune,一种理论严谨、基于策略的微调框架,将GaussMark信号作为奖励,同时正则化防止生成质量退化。结果表明,MarkTune在保持生成质量的同时,持续改善了质量-可检测性权衡,其性能逼近推理时水印水平,且对重述和微调攻击具有鲁棒性,跨数据集泛化能力强。实验验证了该方法在多个数据集上的有效性,确立其为开放权重大模型植入高质量、强鲁棒水印的通用方案。

原文摘要 · Abstract (English)

Watermarking aims to embed hidden signals in generated text that can be reliably detected when given access to a secret key. Open-weight language models pose acute challenges for such watermarking schemes because the inference-time interventions that dominate contemporary approaches cannot be enforced once model weights are public. Existing watermaking techniques for open-weight models, such as the recently proposed GaussMark, typically rely on small modifications to model weights, which can yield signals detectable to those equipped with a secret key, but achieving detection power comparable to inference-time watermarks generally requires weight perturbations that noticeably reduce generation quality. We introduce MarkTune, a theoretically principled, on-policy fine-tuning framework that treats the GaussMark signal as a reward while simultaneously regularizing against degradation in text quality. We derive MarkTune as an improvement on GaussMark and demonstrate that MarkTune consistently improves the quality-detectability trade-off over GaussMark by steering finer-grained, watermark-aware weight updates within the model's representation space while preserving generation quality. Empirically, we show that MarkTune pushes the quality-detectability frontier of GaussMark close to that of inference-time watermarking, remains robust to paraphrasing and fine-tuning attacks, and exhibits strong generalization: a model fine-tuned on one dataset retains substantial watermark detection power on unseen datasets. Together, these results establish MarkTune as a general strategy for embedding robust, high-quality watermarks into open-weight LMs.

水印大模型微调质量优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。