arXiv:2507.05053eess.AScs.SD2025-07被引 6

300人头相关传输函数数据集+工具包,助力个性化空间音频研究

The Extended SONICOM HRTF Dataset and Spatial Audio Metrics Toolbox

  • 扩展至300人测量数据,含200人合成HRTF与优化3D扫描
  • 支持机器学习训练,可分析耳部形态对听觉感知的影响
  • 提供可视化工具包,适合音频算法与人因研究者使用

基于耳机的空间音频依赖头相关传输函数(HRTFs)模拟真实声学环境。由于个体解剖结构差异,每个人的HRTF各不相同。本文发布扩展版SONICOM HRTF数据集,总人数达300人,包含部分参与者的年龄、性别等人口统计信息,增强数据代表性。其中200人的HRTF通过Mesh2HRTF算法合成,并配套预处理的3D头部与耳部扫描模型,优化用于HRTF生成。该数据集支持快速迭代优化合成算法,实现大规模数据自动生成;优化扫描模型便于形态学修改,揭示解剖变化对HRTF的影响;更大的样本量提升机器学习方法的有效性。为支持分析,本文还推出空间音频度量工具箱(SAM Toolbox),一个用于高效分析与可视化的Python工具包,提供可定制化功能,服务于高级研究。整体资源为个性化空间音频研究与开发提供完整支持。

原文摘要 · Abstract (English)

Headphone-based spatial audio uses head-related transfer functions (HRTFs) to simulate real-world acoustic environments. HRTFs are unique to everyone, due to personal morphology, shaping how sound waves interact with the body before reaching the eardrums. Here we present the extended SONICOM HRTF dataset which expands on the previous version released in 2023. The total number of measured subjects has now been increased to 300, with demographic information for a subset of the participants, providing context for the dataset's population and relevance. The dataset incorporates synthesised HRTFs for 200 of the 300 subjects, generated using Mesh2HRTF, alongside pre-processed 3D scans of the head and ears, optimised for HRTF synthesis. This rich dataset facilitates rapid and iterative optimisation of HRTF synthesis algorithms, allowing the automatic generation of large data. The optimised scans enable seamless morphological modifications, providing insights into how anatomical changes impact HRTFs, and the larger sample size enhances the effectiveness of machine learning approaches. To support analysis, we also introduce the Spatial Audio Metrics (SAM) Toolbox, a Python package designed for efficient analysis and visualisation of HRTF data, offering customisable tools for advanced research. Together, the extended dataset and toolbox offer a comprehensive resource for advancing personalised spatial audio research and development.

空间音频HRTF3D扫描合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。