选对声源定位的时频特征,比堆模型更有效。
Systematic Evaluation of Time-Frequency Features for Binaural Sound Source Localization
- 用幅度和相位特征组合提升定位精度
- 双特征组合在域内表现优异,多特征泛化更强
- 低复杂度模型也能达到领先性能
本研究系统评估了用于双耳声源定位(SSL)的时频特征设计,重点分析特征选择如何影响模型在不同条件下的表现。我们采用卷积神经网络(CNN)模型,测试了多种基于幅度的特征(如幅值谱图、双耳强度差-ILD)和基于相位的特征(如相位谱图、双耳相位差-IPD)的组合。在包含匹配与不匹配头相关传递函数(HRTFs)的域内和域外数据上进行评估发现,精心设计的特征组合通常优于单纯增加模型复杂度。对于域内定位,如ILD+IPD等双特征组合已足够;而实现跨内容泛化则需结合通道谱图与ILD、IPD的丰富输入。使用最优特征组合,我们的低复杂度CNN模型实现了具有竞争力的性能。研究强调了特征设计在双耳声源定位中的关键作用,并为特定场景与通用场景的定位任务提供了实用指导。
原文摘要 · Abstract (English)
This study presents a systematic evaluation of time-frequency feature design for binaural sound source localization (SSL), focusing on how feature selection influences model performance across diverse conditions. We investigate the performance of a convolutional neural network (CNN) model using various combinations of amplitude-based features (magnitude spectrogram, interaural level difference - ILD) and phase-based features (phase spectrogram, interaural phase difference - IPD). Evaluations on in-domain and out-of-domain data with mismatched head-related transfer functions (HRTFs) reveal that carefully chosen feature combinations often outperform increases in model complexity. While two-feature sets such as ILD + IPD are sufficient for in-domain SSL, generalization to diverse content requires richer inputs combining channel spectrograms with both ILD and IPD. Using the optimal feature sets, our low-complexity CNN model achieves competitive performance. Our findings underscore the importance of feature design in binaural SSL and provide practical guidance for both domain-specific and general-purpose localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。