arXiv:2606.15187eess.AScs.SD2026-06中稿 · Interspeech 2026

构建音频水印检测基准,评估多种水印在真实环境下的稳定性。

VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations

论文配图:VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations
图 1 · 摘自论文原文
  • 统一注入10种水印方法,覆盖神经与传统算法
  • 引入三类扰动模拟真实录音传输场景,验证检测鲁棒性
  • 提出可扩展的AudioWMD检测器,适合多方法跨分布测试

随着语音生成系统在开放环境中的快速部署,为音频内容提供可验证的来源归属和版权责任变得至关重要。当前研究的一个关键空白是缺乏一个统一的基准,用于系统比较不同水印注入方法在现实分布偏移下的表现。为此,我们构建了VoxWatermark基准,基于多语言、多源语料库,对10种水印方法(4种神经网络、6种传统方法)进行统一注入与标注,并引入无框、黑盒和白盒扰动,以模拟真实的录音与传输条件。基于该基准,我们提出AudioWMD作为大规模、多方法、跨分布场景下的稳健基线检测器。实验结果表明,水印注入方法的多样性及分布偏移会影响检测稳定性,同时验证了AudioWMD的有效性与可扩展性。数据集与代码已公开。

原文摘要 · Abstract (English)

With the rapid deployment of speech generation systems in open environments, providing verifiable source attribution and copyright accountability for audio content has become critical. A gap in current research is the lack of a unified benchmark that systematically compares different watermark injection methods under realistic distribution shifts. To address this, we build VoxWatermark by applying 10 watermarking methods (4 neural and 6 traditional) with unified injection and annotation on multilingual, multi-source corpora, and introducing no-box, black-box, and white-box perturbations to simulate real recording and transmission conditions. Based on this benchmark, we propose AudioWMD as a robust baseline detector for large-scale, multi-method, cross-distribution settings. Results show that injection-method diversity and distribution shifts affect detection stability, while validating the effectiveness and scalability of AudioWMD. Dataset and code are publicly available.

音频水印基准测试鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。