arXiv:2601.02688cs.SDcs.AI2026-01

解决远场多人语音识别中声音混叠问题,提升识别准确率。

Multi-channel multi-speaker transformer for speech recognition

  • 设计多通道多说话人变压器,分离并建模各说话人声学特征。
  • 在SMS-WSJ数据集上相对词错误率降低52.2%,优于现有方法。
  • 适合智能会议、车载语音助手等复杂声学环境应用。

随着远程会议和车载语音助手的发展,远场多人语音识别成为研究热点。近期提出的多通道变压器(MCT)展示了Transformer建模远场声学环境的能力,但因说话人之间的干扰,无法有效编码混合音频中每个说话人的高维声学特征。为此,本文提出多通道多说话人变压器(M2Former),用于远场多人语音识别。在SMS-WSJ基准测试中,M2Former在相对词错误率(WER)上分别比神经波束成形、MCT、双路RNN与变换-平均-拼接以及基于多通道深度聚类的端到端系统提升了9.2%、14.3%、24.9%和52.2%。

原文摘要 · Abstract (English)

With the development of teleconferencing and in-vehicle voice assistants, far-field multi-speaker speech recognition has become a hot research topic. Recently, a multi-channel transformer (MCT) has been proposed, which demonstrates the ability of the transformer to model far-field acoustic environments. However, MCT cannot encode high-dimensional acoustic features for each speaker from mixed input audio because of the interference between speakers. Based on these, we propose the multi-channel multi-speaker transformer (M2Former) for far-field multi-speaker ASR in this paper. Experiments on the SMS-WSJ benchmark show that the M2Former outperforms the neural beamformer, MCT, dual-path RNN with transform-average-concatenate and multi-channel deep clustering based end-to-end systems by 9.2%, 14.3%, 24.9%, and 52.2% respectively, in terms of relative word error rate reduction.

语音识别多说话人远场语音Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。