arXiv:2510.08618eess.AScs.CV2025-10中稿 · ACL

用视觉锚定策略让AI听讲稿时先看后听,减少误识别。

VAPO: End-to-end Slide-Enhanced Speech Recognition with Omni-modal Large Language Models

  • 设计‘看-听’分步推理机制,先提取幻灯片语义再转写语音。
  • 在真实数据集上实体识别错误率显著降低,性能达顶尖水平。
  • 专为幻灯片语音识别构建新基准,支持高质量训练与评估。

全模态大语言模型(OLLMs)因其天然的多模态能力,被视为幻灯片增强语音识别的有前景端到端方案。然而我们发现一个根本问题:视觉干扰,即模型倾向于依赖可见文本而非语音信号,导致虚构未被提及的幻灯片内容。为此,我们提出视觉锚定策略优化(VAPO),旨在重塑模型推理过程,遵循人类‘先看后听’的认知链。具体地,设计时间解耦策略:模型首先在<think>块中提取视觉先验作为语义锚点,随后在<answer>块生成转录。该策略通过多目标强化学习优化。此外,我们构建了SlideASR-Bench,一个综合性基准,包含大规模合成语料用于训练和挑战性真实世界测试集用于评估,以解决实体丰富数据稀缺问题。大量实验表明,VAPO有效消除视觉干扰,在SlideASR-Bench及公开数据集上均达到领先性能,显著降低专业领域中的实体识别错误。

原文摘要 · Abstract (English)

Omni-modal large language models (OLLMs) offer a promising end-to-end solution for slide-enhanced speech recognition due to their inherent multimodal capabilities. However, we found a fundamental issue faced by OLLMs: \textit{Visual Interference}, where models show a bias towards visible text over auditory signals, causing them to hallucinate slide content that was never spoken. To address this, we propose Visually-Anchored Policy Optimization (VAPO), which aims to reshape models' inference process to follow the human-like ``Look-then-Listen'' inference chain. Specifically, we design a temporally decoupled policy: the model first extracts visual priors in a <think> block to serve as semantic anchors, then generates the transcription in an <answer> block. The policy is optimized via multi-objective reinforcement learning. Furthermore, we introduce SlideASR-Bench, a comprehensive benchmark designed to address the scarcity of entity-rich data, comprising a large-scale synthetic corpus for training and a challenging real-world test set for evaluation. We conduct extensive evaluations demonstrating that VAPO effectively eliminates visual interference and achieves state-of-the-art performance on SlideASR-Bench and public datasets, significantly reducing entity recognition errors in specialized domains.

语音识别幻灯片增强多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。