用视频信息修正语音识别错误,提升电视剧字幕准确率。
Speech Recognition on TV Series with Video-guided Post-ASR Correction
- 引入多模态模型融合视频上下文,动态修正语音识别结果。
- 在电视剧数据集上,字幕准确率显著优于传统方法。
- 适合需要高精度影视字幕生成的研究与应用者。
自动语音识别(ASR)在深度学习推动下取得显著进展,广泛应用于对话智能、媒体转录和辅助技术。然而,在电视剧等复杂场景中,多重说话人、重叠语音、领域专有术语及长程上下文依赖仍严重影响转录准确性。现有方法未能显式利用视频提供的丰富时空上下文信息。为此,我们提出视频引导的后处理纠错框架(VPC),采用视频-大语言多模态模型(VLMM)捕捉视频上下文并优化ASR输出。在电视剧基准数据集上的评估表明,该方法在复杂多媒体环境中持续提升转录准确率。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems still struggle in complex environments such as TV series, where multiple speakers, overlapping speech, domain-specific terminology, and long-range contextual dependencies pose significant challenges to transcription accuracy. Existing approaches fail to explicitly leverage the rich temporal and contextual information available in the video. To address this limitation, we propose a Video-Guided Post-ASR Correction (VPC) framework that uses a Video-Large Multimodal Model (VLMM) to capture video context and refine ASR outputs. Evaluations on a TV-series benchmark show that our method consistently improves transcription accuracy in complex multimedia environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。