arXiv:2602.22487eess.AScs.SD2026-02中稿 · IEEE Transactions …被引 1

提出并行处理频谱与空间特征的模型,显著提升动态说话人分离效果。

Moving Speaker Separation via Parallel Spectral-Spatial Processing

  • 分频谱与空间两条并行路径,分别建模不同时间尺度特征
  • 在移动说话人场景下SI-SDR提升1.6-2.2 dB,最快移动时仍超13 dB
  • 适合复杂混响、噪声及快速运动下的语音分离任务

动态环境中的多通道语音分离面临挑战,因时变的空间与频谱特征演化速度不同。现有方法多采用串行架构,迫使单一网络同时建模两类特征,引发建模冲突。本文提出双分支并行谱-空(PS2)架构,分别通过并行流处理频谱与空间特征。频谱分支采用基于BLSTM的频率模块、Mamba-based时序模块与自注意力模块建模频谱特性;空间分支使用双向门控循环单元(BGRU)网络处理随源-麦克风几何关系变化的空间特征。两路特征通过交叉注意力融合机制整合,自适应加权贡献。实验表明,PS2在移动说话人场景下相较现有最先进方法,缩放不变信干比(SI-SDR)提升1.6–2.2 dB,且在不同混响时间(RT60)、噪声水平及源运动速度下均保持稳定分离性能。即使源移动迅速,仍保持超过13 dB的SI-SDR提升。该优势在WHAMR!及我们构建的WSJ0-Demand-6ch-Move数据集上一致验证。

原文摘要 · Abstract (English)

Multi-channel speech separation in dynamic environments is challenging as time-varying spatial and spectral features evolve at different temporal scales. Existing methods typically employ sequential architectures, forcing a single network stream to simultaneously model both feature types, creating an inherent modeling conflict. In this paper, we propose a dual-branch parallel spectral-spatial (PS2) architecture that separately processes spectral and spatial features through parallel streams. The spectral branch uses a bi-directional long short-term memory (BLSTM)-based frequency module, a Mamba-based temporal module, and a self-attention module to model spectral features. The spatial branch employs bi-directional gated recurrent unit (BGRU) networks to process spatial features that encode the evolving geometric relationships between sources and microphones. Features from both branches are integrated through a cross-attention fusion mechanism that adaptively weights their contributions. Experimental results demonstrate that the PS2 outperforms existing state-of-the-art (SOTA) methods by 1.6-2.2 dB in scale-invariant signal-to-distortion ratio (SI-SDR) for moving speaker scenarios, with robust separation quality under different reverberation times (RT60), noise levels, and source movement speeds. Even with fast source movements, the proposed model maintains SI-SDR improvements of over 13 dB. These improvements are consistently observed across multiple datasets, including WHAMR! and our generated WSJ0-Demand-6ch-Move dataset.

语音分离并行结构动态场景信号增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。