arXiv:2511.06458cs.SDcs.AI2025-11被引 1

让音频环境迁移带水印,防伪造又保音质。

EchoMark: Perceptual Acoustic Environment Transfer with Watermark-Embedded Room Impulse Response

  • 在隐空间操作,适应不同长短和衰减的混响特征。
  • 音质评分4.22/5,水印识别准确率超99%。
  • 适合需防伪的语音合成、虚拟现实等场景。

声学环境匹配(AEM)旨在将纯净音频迁移到目标声学环境,支持语音配音和沉浸式虚拟现实等应用。直接从混响语音中恢复相似的房间冲击响应(RIR)提供了更便捷灵活的AEM方案。然而,这一能力若被恶意利用,可能导致任意“环境迁移”,如助长高级语音欺骗攻击或破坏录音证据的真实性。为此,我们提出EchoMark,首个基于深度学习的AEM框架,可生成感知相似且嵌入水印的RIR。通过在隐空间操作,有效应对不同持续时间与能量衰减的RIR特性挑战。联合优化重建感知损失与水印检测损失,实现高质量环境迁移与可靠水印恢复。在多个数据集上的实验表明,EchoMark在房间声学参数匹配性能上与当前最佳的FiNS相当。此外,平均意见分(MOS)达4.22/5,水印检测准确率超过99%,误码率(BER)低于0.3%,充分证明其在保持感知质量的同时实现可靠水印嵌入。

原文摘要 · Abstract (English)

Acoustic Environment Matching (AEM) is the task of transferring clean audio into a target acoustic environment, enabling engaging applications such as audio dubbing and auditory immersive virtual reality (VR). Recovering similar room impulse response (RIR) directly from reverberant speech offers more accessible and flexible AEM solution. However, this capability also introduces vulnerabilities of arbitrary ``relocation" if misused by malicious user, such as facilitating advanced voice spoofing attacks or undermining the authenticity of recorded evidence. To address this issue, we propose EchoMark, the first deep learning-based AEM framework that generates perceptually similar RIRs with embedded watermark. Our design tackle the challenges posed by variable RIR characteristics, such as different durations and energy decays, by operating in the latent domain. By jointly optimizing the model with a perceptual loss for RIR reconstruction and a loss for watermark detection, EchoMark achieves both high-quality environment transfer and reliable watermark recovery. Experiments on diverse datasets validate that EchoMark achieves room acoustic parameter matching performance comparable to FiNS, the state-of-the-art RIR estimator. Furthermore, a high Mean Opinion Score (MOS) of 4.22 out of 5, watermark detection accuracy exceeding 99\%, and bit error rates (BER) below 0.3\% collectively demonstrate the effectiveness of EchoMark in preserving perceptual quality while ensuring reliable watermark embedding.

音频迁移水印技术声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。