arXiv:2603.16278eess.ASeess.SP2026-03被引 1

用迭代算法结构提升语音定位精度与鲁棒性

Speakers Localization Using Batch EM In Unfolding Neural Network

  • 将批量EM算法嵌入编码-解码框架,实现可解释的迭代优化
  • 混响环境下定位准确率显著优于传统批量EM方法
  • 适合需要高鲁棒性的语音定位场景,如智能音箱、会议系统

我们提出一种可解释的批量EM展开网络,用于鲁棒的说话人定位。通过将迭代式EM过程嵌入编码器-EM-解码器架构中,该方法缓解了初始化敏感性问题并提升了收敛性能。实验表明,在混响条件下,该方法的定位准确率和鲁棒性均优于经典批量EM方法。

原文摘要 · Abstract (English)

We propose an interpretable Batch-EM Unfolded Network for robust speaker localization. By embedding the iterative EM procedure within an encoder-EM-decoder architecture, the method mitigates initialization sensitivity and improves convergence. Experiments show superior accuracy and robustness over the classical Batch-EM in reverberant conditions.

语音定位神经网络信号处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。