让机器人像人一样听声辨位,用真实耳部响应数据提升定位精度。
Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots
- 用双耳时频特征融合时间差、强度差和高频谱线索,模拟人耳听觉机制。
- 在合成数据集上实现4.4°定位误差与87.1%检测准确率,优于传统方法。
- 适合研究仿生听觉、人形机器人感知或多音源定位的开发者参考。
本文提出双耳声音事件定位与检测(BiSELD)任务,旨在通过双耳音频联合检测并定位多个声音事件,模仿人类的空间听觉机制。为此,我们构建了名为Binaural Set的合成基准数据集,利用实测头相关传输函数(HRTF)和多样声音事件模拟真实听觉场景。为有效处理该任务,提出一种新型输入特征表示——双耳时频特征(BTFF),包含八个通道:左右梅尔谱图、速度图、谱线索图及双耳时间差(ITD)、强度差(ILD)图,覆盖不同频段和空间轴上的空间线索。基于此,设计了CRNN架构的BiSELDnet模型,学习时频模式与基于HRTF的定位线索。实验表明,各子特征均提升性能:速度图增强检测能力,ITD/ILD图实现精确水平定位,谱线索图捕捉垂直空间信息。最终系统达到0.110的SELD误差、87.1%的F-score和4.4°的定位误差,验证了该框架在模拟人类听觉感知方面的有效性。
原文摘要 · Abstract (English)
This paper introduces Binaural Sound Event Localization and Detection (BiSELD), a task that aims to jointly detect and localize multiple sound events using binaural audio, inspired by the spatial hearing mechanism of humans. To support this task, we present a synthetic benchmark dataset, called the Binaural Set, which simulates realistic auditory scenes using measured head-related transfer functions (HRTFs) and diverse sound events. To effectively address the BiSELD task, we propose a new input feature representation called the Binaural Time-Frequency Feature (BTFF), which encodes interaural time difference (ITD), interaural level difference (ILD), and high-frequency spectral cues (SC) from binaural signals. BTFF is composed of eight channels, including left and right mel-spectrograms, velocity-maps, SC-maps, and ITD-/ILD-maps, designed to cover different spatial cues across frequency bands and spatial axes. A CRNN-based model, BiSELDnet, is then developed to learn both spectro-temporal patterns and HRTF-based localization cues from BTFF. Experiments on the Binaural Set show that each BTFF sub-feature enhances task performance: V-map improves detection, ITD-/ILD-maps enable accurate horizontal localization, and SC-map captures vertical spatial cues. The final system achieves a SELD error of 0.110 with 87.1% F-score and 4.4° localization error, demonstrating the effectiveness of the proposed framework in mimicking human-like auditory perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。