用轻量粒子滤波实现动态说话人跟踪,自适应提升语音增强效果。
Self-Steering Deep Non-Linear Spatially Selective Filters for Efficient Extraction of Moving Speakers under Weak Guidance
- 引入粒子滤波实现低复杂度实时跟踪
- 通过反馈机制显著提升定位精度与语音增强性能
- 适合资源受限的实时语音增强场景
针对动态场景中说话人定位与语音增强的挑战,现有基于深度非线性空间选择性滤波的方法在静止说话人上表现优异,但动态情况下需依赖高计算成本的数据驱动追踪算法。为此,本文提出一种基于轻量级粒子滤波的自适应追踪策略,结合因果顺序处理框架,利用空间选择性滤波输出的增强语音信号作为时序反馈,补偿粒子滤波建模能力不足的问题。合成数据集上的评估显示,两者间的自回归交互显著提升了追踪准确率,实现了强增强性能;真实录音的听感测试进一步验证了该自引导流程在主观评价中优于对比方法。
原文摘要 · Abstract (English)
Recent works on deep non-linear spatially selective filters demonstrate exceptional enhancement performance with computationally lightweight architectures for stationary speakers of known directions. However, to maintain this performance in dynamic scenarios, resource-intensive data-driven tracking algorithms become necessary to provide precise spatial guidance conditioned on the initial direction of a target speaker. As this additional computational overhead hinders application in resource-constrained scenarios such as real-time speech enhancement, we present a novel strategy utilizing a low-complexity tracking algorithm in the form of a particle filter instead. Assuming a causal, sequential processing style, we introduce temporal feedback to leverage the enhanced speech signal of the spatially selective filter to compensate for the limited modeling capabilities of the particle filter. Evaluation on a synthetic dataset illustrates how the autoregressive interplay between both algorithms drastically improves tracking accuracy and leads to strong enhancement performance. A listening test with real-world recordings complements these findings by indicating a clear trend towards our proposed self-steering pipeline as preferred choice over comparable methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。