arXiv:2411.18368cs.CLcs.AI2024-11NAACL

用改写文本增强多模态语音识别,提升多种语言对话识别准确率。

AMPS: ASR with Multimodal Paraphrase Supervision

  • 用参考转录的改写句作为额外监督信号训练模型。
  • 在表现差的语句上启用改写目标,使词错误率降低最多5%。
  • 适合多语言对话场景下的语音识别研究与应用。

自然或对话式多语言语音给当前最先进的自动语音识别(ASR)系统带来诸多挑战。本文提出一种新方法AMPS,通过引入基于改写的监督信号,增强多语言多模态ASR系统,以改善包括印地语、马拉地语、马拉雅拉姆语、卡纳达语和尼亚加语在内的多种语言的对话式语音识别性能。训练时,利用参考转录的改写句作为额外监督,并仅对识别效果较差的语句激活该改写目标。结合最先进的多模态模型SeamlessM4T使用AMPS,实现了最高达5%的词错误率(WER)相对降低。我们通过客观和人工评估指标对系统进行了详细分析。

原文摘要 · Abstract (English)

Spontaneous or conversational multilingual speech presents many challenges for state-of-the-art automatic speech recognition (ASR) systems. In this work, we present a new technique AMPS that augments a multilingual multimodal ASR system with paraphrase-based supervision for improved conversational ASR in multiple languages, including Hindi, Marathi, Malayalam, Kannada, and Nyanja. We use paraphrases of the reference transcriptions as additional supervision while training the multimodal ASR model and selectively invoke this paraphrase objective for utterances with poor ASR performance. Using AMPS with a state-of-the-art multimodal model SeamlessM4T, we obtain significant relative reductions in word error rates (WERs) of up to 5%. We present detailed analyses of our system using both objective and human evaluation metrics.

语音识别多模态改写监督多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。