arXiv:2509.17490eess.ASeess.SP2025-09

用轻量U-Net结构提升多动声源定位精度,计算量更低

FUN-SSL: Full-band Layer Followed by U-Net with Narrow-band Layers for Multiple Moving Sound Source Localization

  • 用全频带层+多尺度窄带U-Net替代传统双路径模块
  • 在多个数据集上定位准确率超现有方法,计算量仅为IPDnet的40%
  • 适合资源受限场景下的实时声源定位应用

沿时间与频谱维度的双路径处理在语音处理中表现优异。尽管如FN-SSL和IPDnet等采用双路径结构的声源定位(SSL)模型在多移动声源定位任务中表现突出,但其计算开销较大。本文提出一种新型SSL架构——FUN-SSL,通过引入U-Net实现多尺度窄带处理以降低计算复杂度。该模型将IPDnet中的全窄网络块(由1个全频带LSTM层和1个窄带LSTM层构成)替换为由1个全频带层和一个包含多尺度窄带层的U-Net组成的FUN模块。同时,在每个U-Net内部的跳跃连接基础上,新增了跨FUN模块间的跳跃连接以增强信息传递。实验表明,FUN-SSL在多个基准数据集上的定位性能优于先前方法,且计算复杂度仅为IPDnet的40%。

原文摘要 · Abstract (English)

Dual-path processing along the temporal and spectral dimensions has shown to be effective in various speech processing applications. While the sound source localization (SSL) models utilizing dual-path processing such as the FN-SSL and IPDnet demonstrated impressive performances in localizing multiple moving sources, they require significant amount of computation. In this paper, we propose an architecture for SSL which introduces a U-Net to perform narrow-band processing in multiple resolutions to reduce computational complexity. The proposed model replaces the full-narrow network block in the IPDnet consisting of one full-band LSTM layer along the spectral dimension followed by one narrow-band LSTM layer along the temporal dimension with the FUN block composed of one Full-band layer followed by a U-net with Narrow-band layers in multiple scales. On top of the skip connections within each U-Net, we also introduce the skip connections between FUN blocks to enrich information. Experimental results showed that the proposed FUN-SSL outperformed previously proposed approaches with computational complexity much lower than that of the IPDnet.

声源定位U-Net低延迟双路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。