arXiv:2601.22792eess.AScs.CL2026-01中稿 · IEEE ICASSP 2026

通过声学与语言联合建模,提升多人语音识别的个性化准确率

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

  • 用说话人嵌入提取目标语音,动态调整词汇表进行上下文引导
  • 在双人混合语音中,英文错误率降低至4.7,日文字符错误率降至8.4
  • 适用于多说话人场景下的个性化语音识别,尤其适合跨语言应用

我们提出CALM,一种用于多说话人自动语音识别(ASR)的联合上下文声学-语言建模框架。在个性化AI场景中,声学与语言线索的联合可用性自然推动了目标说话人条件与重叠对话中上下文偏置的融合。CALM通过端到端框架实现这一融合:利用说话人嵌入驱动的目标说话人提取,以及基于动态词汇表的上下文偏置。我们在模拟英语(LibriSpeechMix)和日语(科罗拉多自发日语混合语料库,CSJMix)数据集上评估CALM。在双说话人混合情况下,其在LibriSpeech2Mix上将有偏词错误率(B-WER)从12.7降至4.7,在CSJMix2(eval3)上将有偏字符错误率(B-CER)从16.6降至8.4,验证了跨语言下联合声学-语言建模的有效性。此外,我们在AMI语料库(IHM-mix条件)上报告结果,以验证在标准语音混合数据上的表现。

原文摘要 · Abstract (English)

We present CALM, a joint Contextual Acoustic-Linguistic Modeling framework for multi-speaker automatic speech recognition (ASR). In personalized AI scenarios, the joint availability of acoustic and linguistic cues naturally motivates the integration of target-speaker conditioning with contextual biasing in overlapping conversations. CALM implements this integration in an end-to-end framework through speaker embedding-driven target-speaker extraction and dynamic vocabulary-based contextual biasing. We evaluate CALM on simulated English (LibriSpeechMix) and Japanese (Corpus of Spontaneous Japanese mixtures, CSJMix). On two-speaker mixtures, CALM reduces biased word error rate (B-WER) from 12.7 to 4.7 on LibriSpeech2Mix and biased character error rate (B-CER) from 16.6 to 8.4 on CSJMix2 (eval3), demonstrating the effectiveness of joint acoustic-linguistic modeling across languages. We additionally report results on the AMI corpus (IHM-mix condition) to validate performance on standardized speech mixtures.

语音识别多说话人个性化联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。