用可学习步长优化语音追踪,提升混响环境下的定位精度。
Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking

- 将迭代算法展开为可微分神经网络,自适应调整更新步长。
- 在混响环境下单说话人追踪任务中RMSE低于传统CREM基线。
- 适合需要实时动态声学场景的语音追踪应用。
我们提出一种深度展开的REM网络,用于在轻微混响环境下鲁棒地追踪单个移动说话人。与依赖固定步长衰减策略的经典REM算法不同,该架构通过将迭代过程展开为可微分层,学习自适应更新策略。引入步长网络(Step Size Network),利用FiLM和位置编码(PE)根据时序上下文和收敛状态动态调整递归权重。在混响条件下追踪单说话人的实验结果表明,所提出的展开网络优于采用空间网格搜索将估计质心映射到物理位置的经典CREM基线。在单说话人追踪任务中,该方法实现了更低的均方根误差(RMSE),凸显其在动态声学场景中的潜力。
原文摘要 · Abstract (English)
We propose a deep unfolded REM network for robust tracking of a single moving speaker in mild reverberant environments. Unlike classical REM algorithms, which rely on fixed-step-size decay schedules, the proposed architecture learns an adaptive update policy by unfolding the iterative procedure into differentiable layers. We introduce a Step Size Network that leverages FiLM and PE to dynamically adjust the recursion weights based on temporal context and convergence state. Experimental results for tracking a single speaker under reverberant conditions demonstrate that the proposed unfolded network outperforms the classical CREM baseline, which employs a spatial grid search to map the estimated centroids to physical positions. In the single-speaker tracking task, the proposed method achieves a lower RMSE than the CREM baseline, highlighting its potential for dynamic acoustic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。