TF-Mamba通过时频融合提升声源定位精度。
TF-Mamba: A Time-Frequency Network for Sound Source Localization
- 融合时间与频率特征,用双向Mamba处理多通道音频
- 在模拟和真实数据集上均超越现有先进方法
- 适合语音增强与分离场景的声源定位研究
声源定位(SSL)利用多通道音频数据确定声源位置,常用于提升语音增强与分离效果。提取空间特征对SSL至关重要,尤其在复杂声学环境中。近期新型结构Mamba在多种序列建模任务中表现优异。本文提出基于Mamba的SSL系统TF-Mamba,通过融合时频特征,利用双向Mamba分别处理时间维度与频率维度的信息。实验在模拟与真实数据集上进行,结果表明TF-Mamba显著优于其他先进方法。代码将在后续公开。
原文摘要 · Abstract (English)
Sound source localization (SSL) determines the position of sound sources using multi-channel audio data. It is commonly used to improve speech enhancement and separation. Extracting spatial features is crucial for SSL, especially in challenging acoustic environments. Recently, a novel structure referred to as Mamba demonstrated notable performance across various sequence-based modalities. This study introduces the Mamba for SSL tasks. We consider the Mamba-based model to analyze spatial features from speech signals by fusing both time and frequency features, and we develop an SSL system called TF-Mamba. This system integrates time and frequency fusion, with Bidirectional Mamba managing both time-wise and frequency-wise processing. We conduct the experiments on the simulated and real datasets. Experiments show that TF-Mamba significantly outperforms other advanced methods. The code will be publicly released in due course.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。