用AI分析600+人数据,找出影响听觉方位判断的关键频段。
A Data-Driven Exploration of Elevation Cues in HRTFs: An Explainable AI Perspective Across Multiple Datasets
- 用卷积神经网络+可解释AI分析11个数据集的头相关传输函数。
- 发现多个关键频率带对垂直方向感知有显著影响。
- 适合做音频定位、虚拟现实与听觉模型研究的人参考。
尽管关于头相关传输函数(HRTFs)和频谱线索的研究已十分深入,但精确的垂直方向感知仍是挑战。本文基于来自11个公开数据集的超过600名受试者数据,采用卷积神经网络(CNN)结合可解释人工智能(XAI)技术,系统探究了垂直方向感知的线索。研究测试了多种HRTF预处理方法,重点评估了模型在单数据集内与跨数据集上的泛化能力与可解释性,验证其在不同受试者及测量设置下的鲁棒性。通过类激活映射(CAM)显著性图,识别出可能驱动垂直分类的关键频率带,为理解频谱特征如何影响垂直感知提供了新洞见。本研究拓展了对多条件下的HRTF建模与听觉方位感知机制的理解。
原文摘要 · Abstract (English)
Precise elevation perception in binaural audio remains a challenge, despite extensive research on head-related transfer functions (HRTFs) and spectral cues. While prior studies have advanced our understanding of sound localization cues, the interplay between spectral features and elevation perception is still not fully understood. This paper presents a comprehensive analysis of over 600 subjects from 11 diverse public HRTF datasets, employing a convolutional neural network (CNN) model combined with explainable artificial intelligence (XAI) techniques to investigate elevation cues. In addition to testing various HRTF pre-processing methods, we focus on both within-dataset and inter-dataset generalization and explainability, assessing the model's robustness across different HRTF variations stemming from subjects and measurement setups. By leveraging class activation mapping (CAM) saliency maps, we identify key frequency bands that may contribute to elevation perception, providing deeper insights into the spectral features that drive elevation-specific classification. This study offers new perspectives on HRTF modeling and elevation perception by analyzing diverse datasets and pre-processing techniques, expanding our understanding of these cues across a wide range of conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。