不依赖数据的波束成形提升多说话人语音识别准确率
Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
- 按方向区域划分麦克风信号,预先做波束成形
- 在AMI数据集上降低11%词错误率,提升27%说话人计数准确率
- 无需训练,适合部署在资源受限的实时语音系统
多通道、多说话人场景下的自动语音识别仍受环境噪声、混响和语音重叠影响。本文提出一种数据无关且无需训练的波束成形方法,在应用端到端多通道多说话人语音识别系统前,基于球面极坐标对特定角度区域的信号进行处理。实验表明,使用一组波束成形信号相比原始麦克风信号能显著提升识别性能;增加用于波束成形的信号数量可进一步提高识别准确率,从而更高效利用多通道信号,同时降低语音识别系统的整体输入负载。在AMI会议语料库上的实验显示,该方法相较未使用波束成形的多通道语音识别基线系统,词错误率最高降低11%,说话人计数准确率相对提升27%。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) in multichannel, multi-speaker scenarios remains challenging due to ambient noise, reverberation and overlapping speakers. In this paper, we propose a beamforming approach that processes specific angular sectors based on their spherical polar coordinates before applying an end-to-end multichannel, multi-speaker ASR system. This method is data-independent and training-free. We demonstrate that using a group of beamformed signals improves ASR performance compared to using the same number of raw microphone signals. Moreover, increasing the number of signals used for beamforming further enhances recognition accuracy, leading to a more efficient use of multichannel signals while reducing the overall input load for the ASR system. We conduct experiments on the AMI meeting corpus, where the proposed method reduces word error rate by up to 11% and improves speaker counting accuracy by up to 27% relative compared to a multichannel ASR baseline system that does not exploit beamforming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。