arXiv:2510.27198eess.AScs.SD2025-10

提出基于归一化ℓp范数的麦克风选择方法,提升远场语音识别效果。

Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm

  • 用归一化ℓp范数评估麦克风信号质量,兼顾信噪比与混响特性
  • 在CHiME-8数据集上降低平均词错误率,优于传统信噪比方法
  • 适合需要高鲁棒性的远场语音识别系统使用

引导源分离(GSS)是基于空间分布麦克风的远场自动语音识别(ASR)系统的常用前端。在使用空间分布麦克风时,参考麦克风的选择对输出信号质量和下游ASR性能有显著影响。当前的GSS语音增强通常采用信噪比(SNR)进行参考麦克风选择,虽利于降噪,但可能忽略各麦克风间早期/晚期混响比(ELR)的差异。本文提出两种基于归一化ℓp范数的参考麦克风选择方法:一种仅使用归一化ℓp范数,另一种结合归一化ℓp范数与SNR,以同时考虑信噪比和混响特性差异。在CHiME-8远场ASR系统上的实验表明,所提ℓp范数方法优于基线方法,显著降低了宏平均词错误率。

原文摘要 · Abstract (English)

Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially distributed microphones, the choice of reference microphone may have a large influence on the quality of the output signal and the downstream ASR performance. In GSS-based speech enhancement, reference microphone selection is typically performed using the signal-to-noise ratio (SNR), which is optimal for noise reduction but may neglect differences in early-to-late-reverberant ratio (ELR) across microphones. In this paper, we propose two reference microphone selection methods for GSS-based speech enhancement that are based on the normalized $\ell_p$-norm, either using only the normalized $\ell_p$-norm or combining the normalized $\ell_p$-norm and the SNR to account for both differences in SNR and ELR across microphones. Experimental evaluation using a CHiME-8 distant ASR system shows that the proposed $\ell_p$-norm-based methods outperform the baseline method, reducing the macro-average word error rate.

语音分离麦克风选择远场识别信号处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。