arXiv:2410.00822cs.SDcs.CL2024-10EMNLP被引 3

用图像中的视觉关键词提升语音识别准确率

VHASR: A Multimodal Speech Recognition System With Vision Hotwords

  • 双流架构分别处理语音和图像,再融合输出
  • 在四个数据集上均优于单模态模型,达当前最优
  • 适合需要高精度语音识别的多模态场景

基于图像的多模态自动语音识别(ASR)模型通过引入与音频相关的图像信息来提升识别性能。然而,部分研究认为图像信息对ASR无帮助。本文提出一种新方法,有效利用音频相关图像信息,构建了名为VHASR的多模态语音识别系统,以视觉关键词作为提示增强模型识别能力。系统采用双流架构,先分别处理语音和图像流,再融合输出结果。在Flickr8k、ADE20k、COCO和OpenImages四个数据集上进行评估,实验表明,VHASR能有效利用图像中的关键信息,显著提升语音识别能力,不仅超越单模态ASR模型,还在现有基于图像的多模态ASR中达到最先进水平。

原文摘要 · Abstract (English)

The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that introducing image information to model does not help improving ASR performance. In this paper, we propose a novel approach effectively utilizing audio-related image information and set up VHASR, a multimodal speech recognition system that uses vision as hotwords to strengthen the model's speech recognition capability. Our system utilizes a dual-stream architecture, which firstly transcribes the text on the two streams separately, and then combines the outputs. We evaluate the proposed model on four datasets: Flickr8k, ADE20k, COCO, and OpenImages. The experimental results show that VHASR can effectively utilize key information in images to enhance the model's speech recognition ability. Its performance not only surpasses unimodal ASR, but also achieves SOTA among existing image-based multimodal ASR.

多模态语音识别视觉提示双流模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。