为开源大模型设计抗修改水印,解决可追溯性难题
Towards Watermarking of Open-Source LLMs
- 提出开源模型水印的耐久性要求,支持合并、量化、微调等场景
- 实测现有方法在模型修改后均失效,无法保持水印完整性
- 构建评估基准,推动更鲁棒的开源模型水印研究
尽管闭源大模型的水印技术已成熟并广泛部署,但这些方法不适用于开源模型,因用户可完全控制解码过程。这一场景虽研究不足,却至关重要,尤其在开源模型性能日益提升的背景下。本文首次系统化地提出开源大模型水印的关键需求,包括对模型合并、量化和微调等常见操作的耐久性,并建立具体评估框架。鉴于这些修改普遍存在,耐久性是水印有效的前提。我们调研并评估了现有方法,发现其均不具备耐久性。同时探讨了改进路径,指出当前仍存在的挑战。希望本工作能推动该重要问题的后续进展。
原文摘要 · Abstract (English)
While watermarks for closed LLMs have matured and have been included in large-scale deployments, these methods are not applicable to open-source models, which allow users full control over the decoding process. This setting is understudied yet critical, given the rising performance of open-source models. In this work, we lay the foundation for systematic study of open-source LLM watermarking. For the first time, we explicitly formulate key requirements, including durability against common model modifications such as model merging, quantization, or finetuning, and propose a concrete evaluation setup. Given the prevalence of these modifications, durability is crucial for an open-source watermark to be effective. We survey and evaluate existing methods, showing that they are not durable. We also discuss potential ways to improve their durability and highlight remaining challenges. We hope our work enables future progress on this important problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。