arXiv:2507.16104eess.AS2025-07

解决多设备麦克风异步问题,提升动态会议语音增强效果

Distributed Asynchronous Device Speech Enhancement via Windowed Cross-Attention

  • 用滑动窗口交叉注意力动态对齐异步麦克风特征
  • 在未知时延和时钟漂移下优于传统方法,收敛更快
  • 适合真实会议场景中多设备协同的语音增强任务

越来越多带麦克风的个人设备为动态会议环境中的临时麦克风阵列提供了可能。然而,现有方法大多针对时间同步的麦克风设置,而现实会议中设备间存在时延和时钟漂移。我们发现,主流的变换-平均-拼接(TAC)模块在异步麦克风情况下表现不足。为此,提出一种滑动窗口交叉注意力模块,可动态对齐所有麦克风的特征。该模块对麦克风排列和数量均不变,且易于集成到现有模型中。此外,还设计了多说话人环境下的最优训练目标。在具有未知时延和时钟漂移的多麦克风噪声混响环境下评估,实验结果表明,该方法在iFaSNet和CRUSE模型上均优于TAC,具备更快收敛与更好学习能力,验证了滑动窗口交叉注意力在异步麦克风设置中的有效性。

原文摘要 · Abstract (English)

The increasing number of microphone-equipped personal devices offers great flexibility and potential using them as ad-hoc microphone arrays in dynamic meeting environments. However, most existing approaches are designed for time-synchronized microphone setups, a condition that may not hold in real-world meeting scenarios, where time latency and clock drift vary across devices. Under such conditions, we found transform-average-concatenate (TAC), a popular module for neural multi-microphone processing, insufficient in handling time-asynchronous microphones. In response, we propose a windowed cross-attention module capable of dynamically aligning features between all microphones. This module is invariant to both the permutation and the number of microphones and can be easily integrated into existing models. Furthermore, we propose an optimal training target for multi-talker environments. We evaluated our approach in a multi-microphone noisy reverberant setup with unknown time latency and clock drift of each microphone. Experimental results show that our method outperforms TAC on both iFaSNet and CRUSE models, offering faster convergence and improved learning, demonstrating the efficacy of the windowed cross-attention module for asynchronous microphone setups.

语音增强异步处理多设备

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。