提升声源定位精度与效率,计算成本不到原方法的2%
IPDnet2: an efficient and improved inter-channel phase difference estimation network for sound source localization
- 用oSpatialNet替代LSTM,增强空间特征提取并提升可扩展性
- 引入频时联合压缩机制,计算量不足原方法2%仍保持定位能力
- 模型规模扩大后性能达当前最佳,适合实时应用场景
IPDnet是我们近期提出的实时声源定位网络,采用交替的全带宽和窄带(B)LSTM分别学习全带宽相关性和窄带差分相位-相位差(DP-IPD)特征,表现优异。但独立处理窄带导致计算复杂度高,且LSTM层数受限影响定位精度。本文提出IPDnet2,在保留高性能的同时显著提升效率:采用oSpatialNet作为主干网络以增强空间线索提取并支持更好扩展;设计简单有效的频时池化机制,压缩频率与时间分辨率以降低计算开销,同时不损失定位能力。实验表明,IPDnet2在定位性能接近IPDnet的前提下,计算成本低于其2%;通过扩大模型规模,实现当前最优的声源定位性能,且整体复杂度仍较低。
原文摘要 · Abstract (English)
IPDnet is our recently proposed real-time sound source localization network. It employs alternating full-band and narrow-band (B)LSTMs to learn the full-band correlation and narrow-band extraction of DP-IPD, respectively, which achieves superior performance. However, processing narrow-band independently incurs high computational complexity and the limited scalability of LSTM layers constrains the localization accuracy. In this work, we extend IPDnet to IPDnet2, improving both localization accuracy and efficiency. IPDnet2 adapts the oSpatialNet as the backbone to enhance spatial cues extraction and provide superior scalability. Additionally, a simple yet effective frequency-time pooling mechanism is proposed to compress frequency and time resolutions and thus reduce computational cost, and meanwhile not losing localization capability. Experimental results show that IPDnet2 achieves comparable localization performance with IPDnet while only requiring less than 2\% of its computation cost. Moreover, the proposed network achieves state-of-the-art SSL performance by scaling up the model size while still maintaining relatively low complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。