用不确定性建模提升音视频语音识别在噪声下的鲁棒性
UBG-Net: An Uncertainty-aware Bayesian Gating Network for Robust Audio-Visual Speech Recognition

- 通过贝叶斯门控机制融合音视频特征,引入信号级不确定性
- 在AVCocktail和LRS2上比当前最优模型降低15.2%错误率
- 适合需要高鲁棒性的实际部署场景,如嘈杂环境语音识别
音视频语音识别系统在真实场景中常因信号退化和分布偏移而性能下降。为此,我们提出统一的不确定性建模框架——不确定感知贝叶斯门控网络(UBG-Net)。UBG-Net包含模态不确定性感知贝叶斯融合(MUBF)机制,将信号级似然不确定性注入贝叶斯网络以建模认知不确定性,从而实现对预训练主干特征的鲁棒融合。推理阶段,提出分布不确定性感知分层投票(DUHV),从蒙特卡洛采样中选择文本,优先考虑频率,相同频率时使用推理得分决定。在AVCocktail和LRS2数据集上的实验表明,UBG-Net整体优于现有最先进基线。消融研究证实MUBF和DUHV能有效过滤噪声,提升融合与解码鲁棒性。
原文摘要 · Abstract (English)
Audio-Visual speech recognition systems often degrade in real-world scenarios due to signal corruption and distribution shifts. To address this, we propose a unified uncertainty-modeling framework, namely the uncertainty-aware Bayesian gating network (UBG-Net). UBG-Net features a Modality Uncertainty-aware Bayesian Fusion (MUBF) mechanism that injects signal-level aleatoric uncertainty into a Bayesian network to model epistemic uncertainty, thereby ensuring robust fusion of pre-trained backbone features. For inference, we introduce Distribution Uncertainty-aware Hierarchical Voting (DUHV) to select transcripts from Monte Carlo samples, prioritizing frequency and using inference scores in case of a tie. Experiments on the AVCocktail and LRS2 datasets demonstrate the overall superiority of UBG-Net compared to SOTA baselines. Ablation studies confirm that MUBF and DUHV effectively filter noise, enhancing fusion and decoding robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。