AWARE通过对抗优化实现抗编辑音频水印,无需模拟攻击即可保持高音质和低误码率。
AWARE: Audio Watermarking with Adversarial Resistance to Edits
- 在时频域通过对抗优化嵌入水印,按感知预算动态调整强度
- 检测器聚合比特级证据,对时序错位和截断有强鲁棒性,误码率持续偏低
- 不依赖攻击模拟,适合真实场景下需要高可靠性的音频水印应用
当前基于学习的音频水印方法通常通过扩大训练时的模拟失真集来提升鲁棒性,但这类近似手段范围有限且易过拟合。本文提出AWARE(Audio Watermarking with Adversarial Resistance to Edits),一种不依赖攻击模拟堆栈和手工可微失真的新方法。嵌入过程在时频域通过对抗优化完成,并受制于与水印强度成比例的感知预算。检测采用时间顺序无关的检测器,结合位级读出头(BRH),将时间证据聚合为每比特一个得分,即使在时序错位或剪切情况下仍能可靠解码。实验表明,AWARE在多种音频编辑下均保持高音质与语音可懂度(PESQ/STOI),且误码率(BER)始终很低,常优于代表性先进学习型系统。
原文摘要 · Abstract (English)
Prevailing practice in learning-based audio watermarking is to pursue robustness by expanding the set of simulated distortions during training. However, such surrogates are narrow and prone to overfitting. This paper presents AWARE (Audio Watermarking with Adversarial Resistance to Edits), an alternative approach that avoids reliance on attack-simulation stacks and handcrafted differentiable distortions. Embedding is obtained through adversarial optimization in the time-frequency domain under a level-proportional perceptual budget. Detection employs a time-order-agnostic detector with a Bitwise Readout Head (BRH) that aggregates temporal evidence into one score per watermark bit, enabling reliable watermark decoding even under desynchronization and temporal cuts. Empirically, AWARE attains high audio quality and speech intelligibility (PESQ/STOI) and consistently low BER across various audio edits, often surpassing representative state-of-the-art learning-based systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。