用后处理分类器纠正音视频语音增强中的错误输出
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing

- 引入后处理分类器修正语音归属错误
- 结合排列不变训练使性能显著提升
- 适合高噪声复杂环境下的语音增强应用
本文针对音视频语音增强(AVSE)系统中因视频质量差和训练测试数据不匹配导致的错误语音输出问题,提出一种后处理分类器(PPC)来修正这些错误,确保增强后的语音准确对应目标说话人。通过在训练中采用mixup策略提升分类器鲁棒性。在AVSE-challenge数据集上的实验表明,将PPC融入AVSE模型可显著提升性能,尤其与排列不变训练(PIT)结合时表现最佳。该方法相比基线模型有大幅性能提升,展现出跨模态与架构的广泛应用潜力。
原文摘要 · Abstract (English)
This paper addresses the prevalent issue of incorrect speech output in audio-visual speech enhancement (AVSE) systems, which is often caused by poor video quality and mismatched training and test data. We introduce a post-processing classifier (PPC) to rectify these erroneous outputs, ensuring that the enhanced speech corresponds accurately to the intended speaker. We also adopt a mixup strategy in PPC training to improve its robustness. Experimental results on the AVSE-challenge dataset show that integrating PPC into the AVSE model can significantly improve AVSE performance, and combining PPC with the AVSE model trained with permutation invariant training (PIT) yields the best performance. The proposed method substantially outperforms the baseline model by a large margin. This work highlights the potential for broader applications across various modalities and architectures, providing a promising direction for future research in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。