arXiv:2601.12345eess.AScs.LG2026-01中稿 · IEEE International…被引 2

通过联合自回归机制,提升动态场景中近距说话人分离的鲁棒性。

Adaptive Rotary Steering with Joint Autoregression for Robust Extraction of Closely Moving Speakers in Dynamic Scenarios

  • 引入联合自回归框架,利用语音时频相关性指导旋转波束成形
  • 在合成数据集上显著优于非自回归方法,尤其在说话人交叉时表现更优
  • 适合复杂动态场景下的多说话人分离,如密集对话或移动声源

最近基于安博森尼克(Ambisonics)的深度空间滤波方法在静止多说话人场景中表现出色,通过将声场旋转至目标说话人方向,再进行多通道增强。然而,在说话人移动的动态声学条件下,传统方法难以实现稳定跟踪。本文提出一种基于目标初始方向条件的交错跟踪算法,自动实现旋转波束成形。对于近距离或交叉的说话人,传统空间线索失效,难以有效增强。为此,我们创新性地将处理后的录音作为额外引导信息,嵌入两个算法中,构建联合自回归框架,利用语音的时频相关性解决空间上复杂的说话人构型问题。结果表明,该方法显著提升了对近距说话人的跟踪与增强性能,在合成数据集上持续优于对比的非自回归方法。真实录音验证了其在多重说话人交叉及声源-麦克风距离变化等复杂场景中的有效性。

原文摘要 · Abstract (English)

Latest advances in deep spatial filtering for Ambisonics demonstrate strong performance in stationary multi-speaker scenarios by rotating the sound field toward a target speaker prior to multi-channel enhancement. For applicability in dynamic acoustic conditions with moving speakers, we propose to automate this rotary steering using an interleaved tracking algorithm conditioned on the target's initial direction. However, for nearby or crossing speakers, robust tracking becomes difficult and spatial cues less effective for enhancement. By incorporating the processed recording as additional guide into both algorithms, our novel joint autoregressive framework leverages temporal-spectral correlations of speech to resolve spatially challenging speaker constellations. Consequently, our proposed method significantly improves tracking and enhancement of closely spaced speakers, consistently outperforming comparable non-autoregressive methods on a synthetic dataset. Real-world recordings complement these findings in complex scenarios with multiple speaker crossings and varying speaker-to-array distances.

语音分离自回归模型动态场景波束成形

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。