arXiv:2509.22744eess.AScs.AI2025-09

利用视频字幕提升语音识别准确率,降低错误率20.5%

Index-MSR: A high-efficiency multimodal fusion framework for speech recognition

  • 通过融合视频字幕等文本信息增强语音识别
  • 在多个数据集上实现当前最优准确率,错词率降20.5%
  • 适合需要高精度音画同步的场景,如语音翻译

得益于大规模数据集和基于大模型的架构,自动语音识别(ASR)系统在准确性方面取得了显著进步。然而,在特定领域术语及语义不连贯的短语中,识别性能仍会明显下降。本文提出Index-MSR,一种高效的多模态语音识别框架。其核心是新颖的多模态融合解码器(MFD),可有效将视频中的文本信息(如字幕、演示幻灯片)融入语音识别过程。这种跨模态整合不仅提升了整体识别准确率,还大幅降低了替换错误。在自建字幕数据集与公开的AVSR数据集上的广泛评估表明,Index-MSR达到当前最优准确率,替换错误率降低20.5%。结果证明该方法能高效利用视频中的文本线索提升语音识别精度,展现出在需严格音频-文本同步的应用(如语音翻译)中的强大潜力。

原文摘要 · Abstract (English)

Driven by large scale datasets and LLM based architectures, automatic speech recognition (ASR) systems have achieved remarkable improvements in accuracy. However, challenges persist for domain-specific terminology, and short utterances lacking semantic coherence, where recognition performance often degrades significantly. In this work, we present Index-MSR, an efficient multimodal speech recognition framework. At its core is a novel Multimodal Fusion Decoder (MFD), which effectively incorporates text-related information from videos (e.g., subtitles and presentation slides) into the speech recognition. This cross-modal integration not only enhances overall ASR accuracy but also yields substantial reductions in substitution errors. Extensive evaluations on both an in-house subtitle dataset and a public AVSR dataset demonstrate that Index-MSR achieves sota accuracy, with substitution errors reduced by 20,50%. These results demonstrate that our approach efficiently exploits text-related cues from video to improve speech recognition accuracy, showing strong potential in applications requiring strict audio text synchronization, such as audio translation.

语音识别多模态融合字幕对齐低错误率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。