arXiv:2606.23712eess.SPcs.AI2026-06

用对比学习强化视听融合,提升语音增强效果

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

论文配图:Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
图 1 · 摘自论文原文
  • 在扩散模型中加入视听对比损失,增强视觉信息利用
  • 低信噪比下噪声抑制和语音重建性能显著提升
  • 适合需要高鲁棒性语音增强的场景或系统集成

视听语音增强(AVSE)利用唇动等视觉线索,在嘈杂环境中恢复语音。近期工作提出基于扩散模型的无监督AVSE,通过交叉注意力将视觉特征条件化到语音扩散模型中,作为后验采样框架的数据驱动先验。尽管相比纯音频方法表现更优,但显式强化跨模态对齐在融合中的作用仍不明确。本文提出在扩散训练目标中加入对比音频-视觉损失,以增强视觉信息的使用,同时保持后验采样框架不变。在匹配与非匹配测试数据上的实验表明,该方法在干扰抑制、信号重建和感知质量方面均有持续改进,尤其在低信噪比条件下增益最大。代码已公开于 https://github.com/cexauce/AV-CA-DiffUSE。

原文摘要 · Abstract (English)

Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work introduced diffusion-based unsupervised AVSE, where a speech diffusion model conditioned on visual features via cross-attention is trained and used as a data-driven prior for posterior sampling-based speech enhancement. Despite promising performance over its audio-only counterpart, the impact of explicitly enforcing cross-modal alignment in the fusion remains unclear. In this work, we propose to augment the diffusion training objective with a contrastive audio-visual loss to encourage stronger use of visual information while keeping the posterior sampling framework unchanged. Experiments across matched and mismatched test data show consistent improvements in interference suppression, signal reconstruction, and perceptual quality, with the largest gains at low SNRs. Code is available at https://github.com/ cexauce/AV-CA-DiffUSE

语音增强扩散模型视听融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。