用交叉注意力提升音频水印的鲁棒性与可追溯性
XAttnMark: Learning Robust Audio Watermarking with Cross-Attention
- 通过生成器与检测器共享部分参数,结合交叉注意力机制实现高效信息提取
- 在多种音频变换下仍保持高检测率与准确归属,对抗生成编辑效果显著
- 适合关注生成式AI时代音频版权保护的研究者与开发者
生成式音频合成与编辑技术的快速发展引发了版权侵权、数据溯源和深度伪造音频传播等严重问题。水印技术通过嵌入不可感知但可识别的信号提供主动防护。尽管现有神经网络方法如WavMark和AudioSeal提升了鲁棒性与音质,仍难以兼顾精准检测与可靠归属。本文提出交叉注意力鲁棒音频水印(XATTNMARK),通过生成器与检测器的部分参数共享、交叉注意力机制实现高效消息检索,并引入时序条件模块优化消息分布。此外,设计了符合心理声学特性的时频掩蔽损失,捕捉精细听觉掩蔽效应,增强水印不可感知性。XATTNMARK在检测与归属任务上均达到当前最优性能,对多种音频变换(包括不同强度的生成编辑)具有强鲁棒性。该工作推动了生成式AI时代音频版权保护与真实性验证的发展。
原文摘要 · Abstract (English)
The rapid proliferation of generative audio synthesis and editing technologies has raised serious concerns about copyright infringement, data provenance, and the spread of misinformation via deepfake audio. Watermarking offers a proactive solution by embedding imperceptible yet identifiable and traceable signals into audio content. While recent neural network-based watermarking methods like WavMark and AudioSeal have improved robustness and quality, they struggle to jointly optimize both robust detection and accurate attribution. This paper introduces Cross-Attention Robust Audio Watermark (XATTNMARK), which bridges this gap by leveraging partial parameter sharing between the generator and the detector, a cross-attention mechanism for efficient message retrieval, and a temporal conditioning module for improved message distribution. Additionally, we propose a psychoacoustic-aligned time-frequency (TF) masking loss that captures fine-grained auditory masking effects, improving watermark imperceptibility. XATTNMARK achieves state-of-the-art performance in both detection and attribution, demonstrating superior robustness against a wide range of audio transformations, including challenging generative editing at varying strengths. This work advances audio watermarking for protecting intellectual property and ensuring authenticity in the era of generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。