仅需目标初始方向,实现多移动说话人语音增强的通用解法
Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics
- 将时频处理与空间滤波解耦,用自回归框架保持性能稳定
- 在动态场景中对交叉、密集移动说话人仍能有效分离语音
- 适用于不同阶数的体感音频阵列,泛化性强
线性空间滤波器(波束成形器)在理想参数设定下可实现鲁棒、通用且可解释的语音增强,并具备性能保障。现代波束成形器通常由深度神经网络参数化,但在多移动说话人方向未知的动态场景中性能下降。本文提出一种数据驱动的波束成形流水线,仅需目标说话人初始方向估计。基于高阶体感音频表示,我们证明了神经时频处理可与线性空间处理解耦,从而实现通用且阵列无关的增强效果。通过在帧级因果框架中引入自回归机制,系统在快速说话人运动和长时录音中保持一致性能。合成数据评估显示,在说话人紧密分布且交叉的挑战性条件下仍具鲁棒性。真实办公室会议场景中的实录验证了跨不同体感音频阶数的泛化能力。
原文摘要 · Abstract (English)
Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers are often parameterized by deep neural networks, whose performance degrades in dynamic scenarios with multiple moving speakers of unknown directions. We propose a data-driven beamforming pipeline, which only requires an estimate of the target's initial direction. Building on a higher-order ambisonics representation, we show that neural temporal-spectral processing can be decoupled from linear spatial processing, and thereby achieve generalizable and array-agnostic enhancement. By incorporating autoregression into a frame-wise causal framework, we maintain consistent performance throughout fast speaker motion and long recordings. Evaluation on synthetic data demonstrates robust enhancement under challenging conditions with closely spaced and crossing speakers. Real-world recordings in a dynamic office meeting scenario complement these findings and show generalizability across varying ambisonics orders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。