仅凭初始位置追踪移动说话人,解决动态场景下的空间混淆问题。
Steering Deep Non-Linear Spatially Selective Filters for Weakly Guided Extraction of Moving Speakers in Dynamic Scenarios
- 用初始位置代替持续方向指引,实现弱引导追踪
- 在动态交叉场景中仍能准确提取目标说话人
- 适合无人工持续标注的实时语音追踪应用
近期基于深度非线性空间滤波的说话人提取方法在目标方向已知且静止时表现优异。然而,在空间动态场景中,由于空间特征随时间变化及出现歧义(如移动说话人交叉),挑战显著增加。虽然静态场景下用户可轻松指向目标方向,但手动追踪移动说话人不切实际。本文提出一种仅依赖目标初始位置的弱引导提取方法,以应对空间动态场景。通过引入自研深度追踪算法,并在合成数据集上采用联合训练策略,证明该方法能有效解决空间歧义,甚至优于一个方向不匹配但强引导的提取方法。
原文摘要 · Abstract (English)
Recent speaker extraction methods using deep non-linear spatial filtering perform exceptionally well when the target direction is known and stationary. However, spatially dynamic scenarios are considerably more challenging due to time-varying spatial features and arising ambiguities, e.g. when moving speakers cross. While in a static scenario it may be easy for a user to point to the target's direction, manually tracking a moving speaker is impractical. Instead of relying on accurate time-dependent directional cues, which we refer to as strong guidance, in this paper we propose a weakly guided extraction method solely depending on the target's initial position to cope with spatial dynamic scenarios. By incorporating our own deep tracking algorithm and developing a joint training strategy on a synthetic dataset, we demonstrate the proficiency of our approach in resolving spatial ambiguities and even outperform a mismatched, but strongly guided extraction method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。