提出可量化调节的水印框架,解决生成内容检测与语义失真间的权衡问题。
Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking
- 基于功率校准的统计框架,建立水印参数与检测力、失真度的定量关系
- 实验证明新方法在多模型多数据集上均能稳定找到最优平衡点
- 适合需要可解释、可优化水印策略的研究者与应用开发者
基于对数概率的水印机制是识别大语言模型生成内容的常用方法,但其有效性受检测能力与语义失真之间的根本权衡制约。现有分析对超参数选择指导有限,实际部署仍依赖经验调参。本文构建了一种功率校准的统计框架,明确揭示了水印超参数、检测效能与失真程度之间的定量关系。该刻画将水印设计转化为有指导的优化问题。基于此,我们推导出在约束条件下实现最优权衡的实用参数选择方法。在多个语言模型和数据集上的广泛实验验证了理论,并表明所提框架能一致地识别出帕累托最优点。
原文摘要 · Abstract (English)
Logit-based watermarking is a widely used mechanism for identifying LLM generated content, yet its effectiveness is governed by a fundamental trade-off between detectability and semantic distortion. Existing analyses provide limited guidance for principled hyperparameter selection, leaving practical deployments reliant on heuristic tuning. In this work, we develop a power-calibrated statistical framework that establishes explicit quantitative relationships between watermark hyperparameters, detection power, and distortion. This characterization transforms watermark design into a guided optimization problem. Building on these results, we derive practical parameter selection procedures that achieve optimal tradeoffs under constraints. Extensive experiments across multiple language models and datasets validate the theory and demonstrate that the proposed framework consistently identifies Pareto-optimal points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。