arXiv:2508.04333eess.AScs.SD2025-08被引 1

用八通道双耳特征提升人形机器人声源定位精度,解决仰角难、前后混等问题。

Binaural Sound Event Localization and Detection Neural Network based on HRTF Localization Cues for Humanoid Robots

  • 设计八通道双耳时频特征,融合多维度听觉线索。
  • 在城市噪声下定位准确率显著优于现有最先进模型。
  • 适合需要精准听觉感知的人形机器人应用。

人形机器人需同时估计声源类型与方向以实现环境感知,但传统双通道输入在仰角估计和前后混淆方面表现不佳。本文提出一种基于双耳输入的声源定位与检测神经网络(BiSELDnet),学习时频模式及头相关传输函数(HRTF)定位线索。引入新型八通道双耳时频特征(BTFF),包含左右梅尔谱图、V图、低于1.5 kHz的双耳时间差(ITD)图、高于5 kHz且具前后不对称性的双耳强度差(ILD)图,以及高于5 kHz用于仰角估计的频谱线索(SC)图。在全向、水平及矢状面测试中验证了BTFF的有效性。采用基于高效Trinity模块的BiSELDnet输出各声源类别的时序方向向量,实现同步检测与定位。提出向量激活图(VAM)用于分析网络学习机制,确认其聚焦于N1凹陷频率以实现仰角估计。在城市背景噪声条件下对比评估显示,所提BiSELD模型显著优于现有最先进的双耳输入声源定位与检测模型。

原文摘要 · Abstract (English)

Humanoid robots require simultaneous sound event type and direction estimation for situational awareness, but conventional two-channel input struggles with elevation estimation and front-back confusion. This paper proposes a binaural sound event localization and detection (BiSELD) neural network to address these challenges. BiSELDnet learns time-frequency patterns and head-related transfer function (HRTF) localization cues from binaural input features. A novel eight-channel binaural time-frequency feature (BTFF) is introduced, comprising left/right mel-spectrograms, V-maps, an interaural time difference (ITD) map (below 1.5 kHz), an interaural level difference (ILD) map (above 5 kHz with front-back asymmetry), and spectral cue (SC) maps (above 5 kHz for elevation). The effectiveness of BTFF was confirmed across omnidirectional, horizontal, and median planes. BiSELDnets, particularly one based on the efficient Trinity module, were implemented to output time series of direction vectors for each sound event class, enabling simultaneous detection and localization. Vector activation map (VAM) visualization was proposed to analyze network learning, confirming BiSELDnet's focus on the N1 notch frequency for elevation estimation. Comparative evaluations under urban background noise conditions demonstrated that the proposed BiSELD model significantly outperforms state-of-the-art (SOTA) SELD models with binaural input.

声源定位双耳信号人形机器人深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。