提出DAVE框架,提升真实场景下语音分离的鲁棒性
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
- 分离音频与视觉信号,避免视觉质量差导致性能下降
- 构建21.9万组混合语音数据集,支持复杂环境训练
- 仅在无参考条件下增强,确保不降低原有指标
真实场景下的音视频语音增强仍面临视觉输入不可靠和缺乏大规模真实声学数据的问题。现有方法通常直接融合视觉特征,易受劣化视觉信号影响。本文提出DAVE框架,首先通过组合式声学增强构建大规模训练语料库DAVE-Corpus,包含219,411组混合语音。其次采用渐进式多目标优化策略,同时提升语音分离、可懂度、说话人身份保持和感知质量。进一步设计认证选择性增强链,仅在无参考分区中应用场景路由、GAN去噪和响度归一化,确保参考指标不下降。在真实音视频语音增强挑战赛上的实验表明,DAVE在多种真实混合场景及视觉退化条件下均表现出强鲁棒性。
原文摘要 · Abstract (English)
Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions. Existing approaches usually fuse visual features directly into the separation network, making them vulnerable to degraded visual signals. In this paper, we present DAVE, a decoupled audio-visual enhancement framework for real-world speech separation. Firstly, to address the data scarcity issue, we construct DAVE-Corpus, a large-scale training corpus with 219,411 mixtures generated from public meeting corpora through combinatorial acoustic augmentation. Then, we introduce a progressive multi-objective optimization strategy to jointly improve speech separation, intelligibility, speaker identity preservation, and perceptual quality. We further develop a certified selective enhancement chain that applies scene routing, GAN-based denoising, and loudness normalization only within the no-reference partition, guaranteeing non-degradation of reference-based metrics. Experimental results on the Real-World Audio-Visual Speech Enhancement Challenge demonstrate the robustness of DAVE under both real-world mixed scenarios and visual degradation conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。