用幻灯片提升会议演讲自动转录,让专业术语更准确。
Do Slides Help? Multi-modal Context for Automatic Transcription of Conference Talks
- 结合语音与幻灯片信息,改进语音识别
- 整体词错误率降低34%,专业术语降35%
- 适合学术会议转录与多模态语音研究者
当前最先进的自动语音识别(ASR)系统主要依赖声学信息,忽视多模态上下文。但在科学演讲中,视觉信息对消歧和适应至关重要。本文聚焦于利用演示幻灯片而非仅说话人图像来提升识别效果。首先,构建了一个包含幻灯片的多模态演讲基准数据集,并自动分析领域术语转录情况;其次,提出一种数据增强方法以缓解缺乏带幻灯片标注数据的问题;最后,基于增强数据训练模型,在所有词汇上实现约34%的相对词错误率下降,领域特定术语下降35%,显著优于基线模型。
原文摘要 · Abstract (English)
State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While most work focus on speaker images to handle noise conditions, this work also focuses on integrating presentation slides for the use cases of scientific presentation. In a first step, we create a benchmark for multi-modal presentation including an automatic analysis of transcribing domain-specific terminology. Next, we explore methods for augmenting speech models with multi-modal information. We mitigate the lack of datasets with accompanying slides by a suitable approach of data augmentation. Finally, we train a model using the augmented dataset, resulting in a relative reduction in word error rate of approximately 34%, across all words and 35%, for domain-specific terms compared to the baseline model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。