arXiv:2604.27866eess.AScs.MM2026-04

构建真实场景下音视频语音识别新基准,验证视觉信息在恶劣环境中的关键作用。

LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

  • 基于真实对话数据构建音视频识别基准,支持多场景、多噪声测试
  • 在严重降噪条件下,视觉信息贡献显著提升,模型性能下降更明显
  • 适合研究真实世界中视觉线索如何增强语音识别的团队使用

我们提出 LRS-VoxMM,一个面向真实世界音视频语音识别(AVSR)的基准数据集。该数据集源自 VoxMM——一个包含多样化真实对话、并经人工转录的语料库。我们筛选出适合 AVSR 的样本,并以 LRS 风格进行预处理,可直接用于现有 AVSR 模型训练与评估。相比常用基准,LRS-VoxMM 覆盖更多样化场景与声学条件。我们还发布带加性噪声、混响和带宽限制的退化测试集,用于评估极端声学恶化下的性能。实验表明,LRS-VoxMM 显著难于 LRS3;随着音频质量下降,视觉信息的贡献愈发明显。该基准推动更真实的 AVSR 评估,鼓励研究视觉信息在复杂现实条件下的作用。

原文摘要 · Abstract (English)

We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversations with human-annotated transcriptions. We select AVSR-suitable samples and preprocess them in an LRS-style format for direct use in existing AVSR pipelines. Compared with commonly used benchmarks, LRS-VoxMM covers a more diverse range of scenarios and acoustic conditions. We also release distorted evaluation sets with additive noise, reverberation, and bandwidth limitation to support evaluation under severe acoustic degradation. Experimental results show that LRS-VoxMM is considerably harder than LRS3 and that the contribution of visual information becomes more evident as the audio signal degrades. LRS-VoxMM supports more realistic AVSR benchmarking and encourages further research on the role of visual information in challenging real-world conditions.

音视频识别真实场景多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。