arXiv:2605.25967cs.LGcs.SD2026-05中稿 · ICML被引 2

无需训练,通过优化词表实现强鲁棒音频水印。

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

论文配图:Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio
图 1 · 摘自论文原文
  • 利用社区发现缩减词表,缓解离散化误差对水印影响。
  • 检测率提升数个数量级,且对音频修改天然抗扰。
  • 适合需要无梯度、免微调的生成内容溯源场景。

随着政策逐步跟上生成式AI的能力,水印技术成为内容溯源的关键。自回归模型的推理时水印不适用于连续模态,因离散化不一致。现有方法通过微调模态词表克服此问题,但失去了训练自由的优势。本文基于离散化中的词汇冗余,提出一种强大且鲁棒的合成音频水印方案。我们理论分析了词元错误对水印检测的影响,并通过社区发现获得的精简词表有效缓解该问题。大量实验表明,该无梯度方法使检测率提升数个数量级,同时具备对音频修改的内置鲁棒性。本工作在多媒体词元级水印上达到新基准,其效果源于离散表示学习的本质。

原文摘要 · Abstract (English)

As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autoregressive models are unfit for continuous modalities due to discretization inconsistencies. Existing methods overcome this by finetuning the modality tokenizers, nullifying the watermark's training-free advantage. In this work, motivated by the vocabulary redundancy of discretization, we propose an elegant solution for powerful and robust watermarking of synthetic audio. We theoretically analyze the impact of token errors on watermark detection, and effectively mitigate them using a reduced vocabulary obtained via community detection. Thorough experiments showcase that our gradient-free method can boost detectability by several orders of magnitude, while also achieving built-in robustness to audio modifications. Broadly, we discover a new state-of-the-art for token-level watermarks in multimedia, which simply arises from the nature of discrete representation learning.

音频水印无梯度词表优化生成内容溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。