arXiv:2504.15066cs.MMcs.AI2025-04被引 5

中文语音识别新数据集,结合口型与演讲幻灯片提升识别率

Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides

  • 构建含口型和幻灯片的多模态中文语音数据集
  • 幻灯片信息使识别准确率提升25%,口型提升8%
  • 适合做多模态语音识别或中文ASR研究者使用

将视觉信息融入自动语音识别(ASR)已带来显著提升。然而,现有音视频语音识别(AVSR)数据集与方法通常仅依赖口型或演讲视频上下文,忽视了两者在真实语境中的协同潜力。本文发布一个多模态中文AVSR数据集Chinese-LiPS,包含100小时的语音、视频及人工转录文本,视觉模态涵盖口型与演讲者使用的幻灯片。基于此,我们提出LiPS-AVSR框架,同时利用口型和幻灯片作为视觉输入。实验表明,口型信息可使ASR性能提升约8%,幻灯片信息提升约25%,二者联合提升达35%。数据集已开放:https://kiri0824.github.io/Chinese-LiPS/

原文摘要 · Abstract (English)

Incorporating visual modalities to assist Automatic Speech Recognition (ASR) tasks has led to significant improvements. However, existing Audio-Visual Speech Recognition (AVSR) datasets and methods typically rely solely on lip-reading information or speaking contextual video, neglecting the potential of combining these different valuable visual cues within the speaking context. In this paper, we release a multimodal Chinese AVSR dataset, Chinese-LiPS, comprising 100 hours of speech, video, and corresponding manual transcription, with the visual modality encompassing both lip-reading information and the presentation slides used by the speaker. Based on Chinese-LiPS, we develop a simple yet effective pipeline, LiPS-AVSR, which leverages both lip-reading and presentation slide information as visual modalities for AVSR tasks. Experiments show that lip-reading and presentation slide information improve ASR performance by approximately 8\% and 25\%, respectively, with a combined performance improvement of about 35\%. The dataset is available at https://kiri0824.github.io/Chinese-LiPS/

多模态语音识别中文数据集视觉信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。