系统梳理大模型水印技术,助你选对方案防内容滥用
A Survey on LLM Watermarking: Theory and Deployment
- 按生成/训练时机、检测权限等四维度分类,理清水印设计选择
- 对比不同方法在可检测性、鲁棒性上的权衡,覆盖多种攻击场景
- 适合关注模型可信审计与合规部署的研究者和工程师
大语言模型日益嵌入高影响力工作流,其大规模生成流畅文本的能力加剧了来源模糊、模型误用和内容洗稿风险。模型水印通过在输出中嵌入不可见签名,成为溯源、审计与信任决策的重要技术手段。然而文献快速且不均衡发展:现有分类常混杂独立设计要素,难以比较方法、推断保障或转化为可部署系统。本综述提供面向部署的系统性回顾,围绕从业者需回答的核心问题组织框架:水印嵌入位置(生成时或训练时,词元或表征层)、检测权限(公开或私有)、假设条件(是否可访问logits、采样控制、密钥、模型所有权),以及目标威胁模型(改写、翻译、摘要、风格迁移、词元篡改、自适应移除)。我们梳理主流技术——采样偏置、基于编码、表征与训练驱动方法,并从可检测性、鲁棒性、分布偏移角度分析其安全-效用权衡。进一步综述攻击与逃避策略、评估协议与指标(如误报控制、校准、鲁棒性曲线),并指出跨模型迁移、多模态流程、合谋攻击与治理约束等开放挑战。最后,为真实业务需求提供水印设计选型指南,并提出实现可靠、可问责大模型部署所需的研究方向。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly embedded in high-impact workflows, yet their ability to generate fluent text at scale has amplified risks of provenance ambiguity, model misuse, and large-scale content laundering. LLM watermarking, embedding invisible signatures into model outputs, has emerged as a promising technical layer for attribution, auditing, and downstream trust decisions. However, the literature has grown rapidly and unevenly: existing categorizations often mix orthogonal design choices, making it difficult to compare methods, reason about guarantees, or translate research results into deployable systems. This survey provides a systematic, deployment-oriented review of LLM watermarking. We organize the space by the core questions practitioners must answer: where a watermark is embedded (generation-time vs. training-time, token vs. representation), who can detect it (public vs. private detection authority), what is assumed (access to logits, sampling control, secret keys, model ownership), and which threat models are targeted (paraphrasing, translation, summarization, style transfer, token manipulation, and adaptive removal). We synthesize the main families of techniques-including sampling biasing, code-based schemes, representation- and training-based approaches-and analyze their security-utility trade-offs through the lens of detectability, robustness, and distribution shift. We further review attack and evasion strategies, evaluation protocols and metrics (false positive control, calibration, robustness curves), and open challenges such as cross-model transfer, multi-modal pipelines, collusion, and governance constraints. Finally, we provide practical guidance for selecting watermark designs under real operational requirements and identify research directions needed for reliable, accountable LLM deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。