arXiv:2506.13455eess.AScs.SD2025-06

用BiMamba替代Transformer,提升立体声事件定位精度并降低计算开销

Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling

  • 用BiMamba替代Conformer,更好建模时空关系
  • 在DCASE2025数据集上性能优于基线和原版PSELDnet
  • 计算复杂度更低,适合资源受限场景部署

预训练方法在声音事件定位与检测(SELD)任务中取得显著进展,但现有基于Transformer的模型存在计算复杂度高的问题。本文提出一种基于预训练PSELDnet与双向Mamba序列建模的立体声SELD系统。将Conformer模块替换为BiMamba模块,并引入非对称卷积,以更有效地建模时间与频率维度间的时空关系。实验结果表明,所提方法在DCASE2025 Task 3开发集上显著优于基线模型及原始PSELDnet(采用Conformer解码器架构),同时降低了计算复杂度。这些发现验证了BiMamba架构在解决SELD任务挑战中的有效性。

原文摘要 · Abstract (English)

Pre-training methods have achieved significant performance improvements in sound event localization and detection (SELD) tasks, but existing Transformer-based models suffer from high computational complexity. In this work, we propose a stereo sound event localization and detection system based on pre-trained PSELDnet and bidirectional Mamba sequence modeling. We replace the Conformer module with a BiMamba module and introduce asymmetric convolutions to more effectively model the spatiotemporal relationships between time and frequency dimensions. Experimental results demonstrate that the proposed method achieves significantly better performance than the baseline and the original PSELDnet with Conformer decoder architecture on the DCASE2025 Task 3 development dataset, while also reducing computational complexity. These findings highlight the effectiveness of the BiMamba architecture in addressing the challenges of the SELD task.

声音定位Mamba音频处理预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。