arXiv:2410.04478cs.SDcs.CL2024-10

通过语音摘要向量实现可配置多语言语音识别,提升准确率与灵活性。

Configurable Multilingual ASR with Speech Summary Representations

  • 引入语音摘要向量与适配器,按语句级融合多语言模块输出
  • 在7种语言的MLS数据集上将错误率降至9.95%,优于基线
  • 支持手动提示或自动适配,适合跨语言场景应用

全球约一半人口为多语言使用者,多语言语音识别(MASR)因此至关重要。当目标语言未知时,部署多个单语言模型面临挑战。为此,本文提出可配置多语言语音识别模型csvMASR,其创新性地引入语音摘要向量表示,借鉴对话分离中的摘要机制,在语句层级整合各语言专用组件输出,并结合辅助语言分类损失增强可配置性。基于多语言LibriSpeech(MLS)数据集中的7种语言实验表明,csvMASR性能优于现有模型,词错误率(WER)从10.33%降至9.95%。同时,该模型在语言分类与提示任务中表现更优。

原文摘要 · Abstract (English)

Approximately half of the world's population is multilingual, making multilingual ASR (MASR) essential. Deploying multiple monolingual models is challenging when the ground-truth language is unknown in advance. This motivates research efforts on configurable multilingual MASR models that can be prompted manually or adapted automatically to recognise specific languages. In this paper, we present the Configurable MASR model with Summary Vector (csvMASR), a novel architecture designed to enhance configurability. Our approach leverages adapters and introduces speech summary vector representations, inspired by conversational summary representations in speech diarization, to combine outputs from language-specific components at the utterance level. We also incorporate an auxiliary language classification loss to enhance configurability. Using data from 7 languages in the Multilingual Librispeech (MLS) dataset, csvMASR outperforms existing MASR models and reduces the word error rate (WER) from 10.33\% to 9.95\% when compared with the baseline. Additionally, csvMASR demonstrates superior performance in language classification and prompting tasks.

多语言识别语音摘要可配置模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。