测试音频水印在神经压缩等真实干扰下的鲁棒性,发现现有方法普遍脆弱。
A Comprehensive Real-World Assessment of Audio Watermarking Algorithms: Will They Survive Neural Codecs?
- 构建真实场景攻击管道,涵盖压缩、混响等多重干扰
- 神经压缩是最大威胁,训练时包含压缩也难完全抵御
- 极性反转、时间拉伸等特定干扰会严重破坏部分方法
我们提出鲁棒音频水印基准(RAW-Bench),用于系统评估基于深度学习的音频水印方法。为模拟真实使用场景,引入包含压缩、背景噪声、混响等多种失真的综合攻击流程,并采用涵盖语音、环境音和音乐的多样化测试数据集。在RAW-Bench上评估四种现有水印方法发现:(i) 神经压缩技术构成最严峻挑战,即使算法训练时包含此类压缩;(ii) 通常通过对抗攻击训练可提升鲁棒性,但某些情况下仍不足。此外,极性反转、时间拉伸或混响等特定失真对部分方法造成严重影响。评估框架已开源:github.com/SonyResearch/raw_bench。
原文摘要 · Abstract (English)
We introduce the Robust Audio Watermarking Benchmark (RAW-Bench), a benchmark for evaluating deep learning-based audio watermarking methods with standardized and systematic comparisons. To simulate real-world usage, we introduce a comprehensive audio attack pipeline with various distortions such as compression, background noise, and reverberation, along with a diverse test dataset including speech, environmental sounds, and music recordings. Evaluating four existing watermarking methods on RAW-bench reveals two main insights: (i) neural compression techniques pose the most significant challenge, even when algorithms are trained with such compressions; and (ii) training with audio attacks generally improves robustness, although it is insufficient in some cases. Furthermore, we find that specific distortions, such as polarity inversion, time stretching, or reverb, seriously affect certain methods. The evaluation framework is accessible at github.com/SonyResearch/raw_bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。