arXiv:2606.07259eess.AScs.SD2026-06中稿 · Interspeech 2026 L…

构建精准匹配的测试集,发现当前音视频语音识别模型泛化能力严重不足。

Assessing True Generalisability of Audio-Visual Speech Recognisers

  • 从MultiVSR中提取与LRS3分布完全一致的测试集,严格控制音视频与人群特征
  • 五种顶尖模型在该测试集上性能骤降,证明存在严重过拟合
  • 揭示词汇偏差和错误模式,且音视频融合反而弱于纯音频方案

当前音视频语音识别(AVSR)模型在标准LRS3基准上表现接近完美,引发对适应性过拟合的担忧。为系统评估真实泛化能力,我们从海量MultiVSR数据集中构建了一个高度受控、未见过的评估子集,其声学、视觉及人口统计分布与LRS3测试集严格匹配。评估五种顶尖架构发现,性能普遍崩溃,证明现有系统即使在条件严格对齐下仍无法泛化。通过七项属性的细粒度分析,我们确定了性能退化的具体驱动因素。此外,我们发现显著的词汇偏差,揭示了不同的错误模式,并意外发现音视频融合表现甚至落后于纯音频设置。我们公开了该匹配测试集,供未来基准测试使用。

原文摘要 · Abstract (English)

Current Audio-Visual Speech Recognition (AVSR) models achieve near-perfect performance on the standard LRS3 benchmark, raising concerns of adaptive overfitting. To systematically assess true generalisability, we construct a highly controlled, unseen evaluation set subsampled from the massive MultiVSR dataset. Unlike standard out-of-distribution benchmarks, our subset strictly matches the acoustic, visual, and demographic distributions of the LRS3 test set. Evaluating five state-of-the-art architectures reveals a universal performance collapse, proving that current systems fail to generalise even under strictly aligned conditions. Through a fine-grained attribute analysis across seven factors, we isolate the specific drivers of this degradation. Furthermore, we uncover a profound lexical bias, expose distinct error patterns, and surprisingly reveal that audio-visual performance even lags behind audio-only settings. We release our matched test set for future benchmarking.

音视频识别泛化能力基准测试过拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。