arXiv:2605.18442eess.AS2026-05

让语音分离模型适应不同麦克风布局,提升泛化能力。

Flexible Multi-Channel Target Speaker Extraction Using Geometry-Conditioned Spatially Selective Non-linear Filters

论文配图:Flexible Multi-Channel Target Speaker Extraction Using Geometry-Conditioned Spatially Selective Non-linear Filters
图 1 · 摘自论文原文
  • 用几何条件控制滤波器,动态调整空间特征提取
  • 在三种阵列上测试均表现更好,跨布局泛化强
  • 适合部署在多变硬件环境的语音增强系统

近期提出的空间选择性非线性滤波器(SSF)利用目标说话人方向(DOA)作为空间线索进行语音分离,但其性能在阵列几何不匹配时显著下降。本文提出几何条件化SSF(GC-SSF),引入基于FiLM层的几何条件分支,并设计联合编码DOA与麦克风位置的特征(DOA-MPE)。该条件分支通过DOA-MPE调制SSF的中间特征图,以捕捉麦克风布局与目标说话人之间的空间关系。在圆形、均匀线性和随机麦克风阵列上的实验表明,所提方法在不匹配几何下仍保持优异的泛化能力和高空间选择性,有效适配不同阵列配置的语音分离任务。

原文摘要 · Abstract (English)

Recently, a spatially selective non-linear filter (SSF) has been proposed for target speaker extraction, using the target direction-of-arrival (DOA) as a spatial cue. Since learned intermediate features are tied to the microphone geometry, the performance of the SSF degrades significantly when evaluated on mismatched array geometries. In this paper, we propose a geometry-conditioned SSF (GC-SSF), which incorporates a geometry-conditioning branch based on FiLM layers. Furthermore, we propose a feature that jointly encodes the DOA and the microphone positions (DOA-MPE). The conditioning branch modulates the intermediate feature maps of the SSF using the DOA-MPE feature to capture the spatial relationship between the microphone positions and the target speaker. Experimental results across circular, uniform linear, and random microphone arrays show that the proposed GC-SSF generalizes better to mismatched geometries while maintaining high spatial selectivity, demonstrating its ability to effectively adapt the filtering process to different array geometries

语音分离麦克风阵列空间建模泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。