arXiv:2601.18393cs.SDcs.CL2026-01

图文联合模型提升中英文语音识别,准确率显著优于现有方法

OCR-Enhanced Multimodal ASR Can Read While Listening

  • 融合线性与Q-Former结构,通过交叉注意力生成更强视听特征
  • 英语词错误率降5.75%,中文字符错误率降16.5%(相比Whisper)
  • 适用于需要跨语言视听识别的场景,如影视字幕生成

视觉信息(如电影字幕)常有助于自动语音识别。本文提出Donut-Whisper,一种双编码器的音视频语音识别模型,利用视觉信息提升英、中文语音识别性能。该模型通过交叉注意力模块结合线性结构与基于Q-Former的模态对齐优势,生成更强大的音视频特征。同时,提出轻量级知识蒸馏方案,使音视频模型可指导纯音频模型提升性能。此外,构建了一个包含中英文片段的新型多语言音视频语音识别数据集。实验表明,Donut-Whisper在该数据集的英、中文分区上均显著优于Donut和Whisper large V3基线模型,分别实现5.75%绝对词错误率(WER)降低和16.5%绝对字符错误率(CER)降低。

原文摘要 · Abstract (English)

Visual information, such as subtitles in a movie, often helps automatic speech recognition. In this paper, we propose Donut-Whisper, an audio-visual ASR model with dual encoder to leverage visual information to improve speech recognition performance in both English and Chinese. Donut-Whisper combines the advantage of the linear and the Q-Former-based modality alignment structures via a cross-attention module, generating more powerful audio-visual features. Meanwhile, we propose a lightweight knowledge distillation scheme showcasing the potential of using audio-visual models to teach audio-only models to achieve better performance. Moreover, we propose a new multilingual audio-visual speech recognition dataset based on movie clips containing both Chinese and English partitions. As a result, Donut-Whisper achieved significantly better performance on both English and Chinese partition of the dataset compared to both Donut and Whisper large V3 baselines. In particular, an absolute 5.75% WER reduction and a 16.5% absolute CER reduction were achieved on the English and Chinese sets respectively compared to the Whisper ASR baseline.

语音识别图文融合多语言模型蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。