arXiv:2603.23723eess.AScs.LG2026-03被引 1

用自回归反馈提升动态说话人追踪,实时高效且效果更好

Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers

  • 通过自回归方式将增强语音反馈给贝叶斯追踪器,实现动态优化
  • 在真实场景下提升追踪精度,计算开销几乎不变
  • 适用于需要实时处理的移动说话人分离任务

深度空间选择性滤波器在已知方向的静止说话人场景中可实现高质量增强并具备实时能力。为在仅知初始方向的动态场景中保持同等性能,需准确且轻量级的追踪算法。假设逐帧因果处理,时间反馈允许利用增强后的语音信号改进追踪表现。本文研究如何将增强信号融入轻量级追踪算法,并自回归引导深度空间滤波器。所提出的贝叶斯追踪方法兼容任意深度空间滤波器。为提升模拟轨迹的真实性,开发基于社会力模型的合成数据生成框架。结果验证,自回归融合显著提升贝叶斯追踪器精度,带来更优增强效果,计算开销几乎不增加。真实录音进一步证实方法对未见声学条件的泛化能力。

原文摘要 · Abstract (English)

Deep spatially selective filters achieve high-quality enhancement with real-time capable architectures for stationary speakers of known directions. To retain this level of performance in dynamic scenarios where only the speakers' initial directions are given, accurate, yet computationally lightweight tracking algorithms become necessary. Assuming a frame-wise causal processing style, temporal feedback allows for leveraging the enhanced speech signal to improve tracking performance. In this work, we investigate strategies to incorporate the enhanced signal into lightweight tracking algorithms and autoregressively guide deep spatial filters. Our proposed Bayesian tracking algorithms are compatible with arbitrary deep spatial filters. To increase the realism of simulated trajectories during development and evaluation, we develop a synthetic data generation framework based on the social force model. Results validate that the autoregressive incorporation significantly improves the accuracy of our Bayesian trackers, resulting in superior enhancement with none or only negligibly increased computational overhead. Real-world recordings complement these findings and demonstrate the generalizability of our methods to unseen acoustic conditions.

语音增强动态追踪自回归贝叶斯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。