首个针对视频会议的多模态语音识别数据集,揭示了性能下降的关键原因。
When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
- 构建首个专用于视频会议的多模态数据集MLD-VC,包含31人、22.79小时数据
- 发现语音增强算法导致频谱分布偏移,是性能下降主因,使第一/二共振峰改变
- 在真实视频会议平台上,微调模型可降低平均17.5%的词错误率
音频-视觉语音识别(AVSR)在离线场景取得显著进展,但在真实视频会议(VC)环境中的鲁棒性仍缺乏研究。本文首次系统评估主流VC平台上的先进AVSR模型,揭示传输失真与人类自发夸张表达导致严重性能下降。为此,我们构建了首个面向视频会议的多模态数据集MLD-VC,包含31名说话人、22.79小时音视频数据,并明确引入隆巴德效应以增强人类夸张表达。通过全面分析发现,语音增强算法是分布偏移的主要来源,改变了音频的第一和第二共振峰。有趣的是,隆巴德效应引起的分布偏移与语音增强高度相似,解释了为何在隆巴德数据上训练的模型在视频会议中更具鲁棒性。对MLD-VC进行微调可显著缓解该问题,在多个视频会议平台上实现平均17.5%的词错误率(CER)降低。研究成果与数据集为开发更鲁棒、泛化性强的现实视频会议AVSR系统奠定基础。MLD-VC已公开于https://huggingface.co/datasets/nccm2p2/MLD-VC。
原文摘要 · Abstract (English)
Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored. This paper presents the first systematic evaluation of state-of-the-art AVSR models across mainstream VC platforms, revealing severe performance degradation caused by transmission distortions and spontaneous human hyper-expression. To address this gap, we construct \textbf{MLD-VC}, the first multimodal dataset tailored for VC, comprising 31 speakers, 22.79 hours of audio-visual data, and explicit use of the Lombard effect to enhance human hyper-expression. Through comprehensive analysis, we find that speech enhancement algorithms are the primary source of distribution shift, which alters the first and second formants of audio. Interestingly, we find that the distribution shift induced by the Lombard effect closely resembles that introduced by speech enhancement, which explains why models trained on Lombard data exhibit greater robustness in VC. Fine-tuning AVSR models on MLD-VC mitigates this issue, achieving an average 17.5% reduction in CER across several VC platforms. Our findings and dataset provide a foundation for developing more robust and generalizable AVSR systems in real-world video conferencing. MLD-VC is available at https://huggingface.co/datasets/nccm2p2/MLD-VC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。