让大模型输出只能被指定检测器识别,且不影响正常使用。
Multi-Designated Detector Watermarking for Language Models
- 用多指定验证签名技术实现定向水印,仅特定检测器可识别。
- 水印不影响输出质量,且支持所有权声明功能。
- 适合需控制模型输出归属的厂商或机构使用。
本文首次提出针对大语言模型的多指定检测器水印(MDDW)机制。该技术使模型提供方可生成仅由特定(可能多个)指定检测器识别的水印输出,同时保证普通用户无感知的质量下降。我们为MDDW形式化定义了安全标准,并提出基于多指定验证签名(MDVS)的通用构造框架。鉴于大模型输出的高经济价值,引入可声明性作为可选安全特性,使提供方可在指定检测器环境下主张输出所有权。为此,我们提出一种通用转换方法,将任意MDVS转化为可声明式MDVS。实现表明该方案具备先进功能与灵活性,性能指标优异。
原文摘要 · Abstract (English)
In this paper, we initiate the study of \emph{multi-designated detector watermarking (MDDW)} for large language models (LLMs). This technique allows model providers to generate watermarked outputs from LLMs with two key properties: (i) only specific, possibly multiple, designated detectors can identify the watermarks, and (ii) there is no perceptible degradation in the output quality for ordinary users. We formalize the security definitions for MDDW and present a framework for constructing MDDW for any LLM using multi-designated verifier signatures (MDVS). Recognizing the significant economic value of LLM outputs, we introduce claimability as an optional security feature for MDDW, enabling model providers to assert ownership of LLM outputs within designated-detector settings. To support claimable MDDW, we propose a generic transformation converting any MDVS to a claimable MDVS. Our implementation of the MDDW scheme highlights its advanced functionalities and flexibility over existing methods, with satisfactory performance metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。