融合状态空间与自注意力,提升声源定位精度与效率
State Space and Self-Attention Collaborative Network with Feature Aggregation for DOA Estimation
- 通过特征聚合增强时频维度信息,提升定位基础
- 轻量化结构使模型在保持高精度下计算量更低
- 适合需要实时处理的声源定位场景
由于声学特性在时频域持续变化,声源到达方向(DOA)估计面临挑战。准确定位依赖于有效聚合相关特征并建模时间依赖性。在时序建模中,性能与效率的平衡仍是难题。为此,我们提出FA-Stateformer:一种结合状态空间与自注意力的协同网络,包含特征聚合模块,用于增强时频维度的有用特征;采用受挤压-激励机制启发的轻量级Conformer架构,压缩前馈层以减少冗余和参数开销;引入时间移位机制,在保持小卷积核的前提下扩大感受野;并加入双向Mamba模块,实现前后向高效的状态空间建模。剩余自注意力层与Mamba块协同工作,形成兼顾表达能力与计算效率的框架。大量实验表明,相比传统架构,该模型在性能与效率上均表现更优。
原文摘要 · Abstract (English)
Accurate direction-of-arrival (DOA) estimation for sound sources is challenging due to the continuous changes in acoustic characteristics across time and frequency. In such scenarios, accurate localization relies on the ability to aggregate relevant features and model temporal dependencies effectively. In time series modeling, achieving a balance between model performance and computational efficiency remains a significant challenge. To address this, we propose FA-Stateformer, a state space and self-attention collaborative network with feature aggregation. The proposed network first employs a feature aggregation module to enhance informative features across both temporal and spectral dimensions. This is followed by a lightweight Conformer architecture inspired by the squeeze-and-excitation mechanism, where the feedforward layers are compressed to reduce redundancy and parameter overhead. Additionally, a temporal shift mechanism is incorporated to expand the receptive field of convolutional layers while maintaining a compact kernel size. To further enhance sequence modeling capabilities, a bidirectional Mamba module is introduced, enabling efficient state-space-based representation of temporal dependencies in both forward and backward directions. The remaining self-attention layers are combined with the Mamba blocks, forming a collaborative modeling framework that achieves a balance between representation capacity and computational efficiency. Extensive experiments demonstrate that FA-Stateformer achieves superior performance and efficiency compared to conventional architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。