arXiv:2603.24104eess.AS2026-03被引 2

用手机拍头像就能生成人耳声学模型,但精度不够高。

Photogrammetry-Reconstructed 3D Head Meshes for Accessible Individual Head-Related Transfer Functions

  • 用苹果拍摄工具重建头部3D模型,自动合成个体声学参数。
  • 相比实测数据,方向识别错误率高出三倍以上。
  • 适合对声学精度要求不高的普通用户或快速原型设计。

个体化头相关传输函数(HRTFs)对精确的空间音频双耳渲染至关重要,但因测量复杂而难以获取。本研究探讨使用消费级硬件通过摄影测量法重建的头耳三维网格(PR网格)是否可作为个体化HRTF合成的实用基准。基于SONICOM HRTF数据集,为150名受试者采集了72张图像,并利用Apple的Object Capture API生成PR网格。采用Mesh2HRTF方法计算出对应的合成HRTFs,与实测HRTFs、高分辨率扫描导出的HRTFs、KEMAR及随机HRTFs在数值评估、听觉模型和行为定位实验(N=27)中进行对比。结果表明,PR合成的HRTFs保留了到达时间差(ITD)线索,但在双耳强度差(ILD)和频谱特性上存在显著误差。听觉模型预测与行为数据均显示其方位误差率显著升高、仰角分辨能力下降,前后混淆现象更严重,感知性能甚至低于随机HRTFs。当前摄影测量流程虽支持个体化HRTF生成,但受限于耳廓形态细节不足及高频频谱保真度欠缺,无法充分还原含单耳线索的精确个体化声学特征。

原文摘要 · Abstract (English)

Individual head-related transfer functions (HRTFs) are essential for accurate spatial audio binaural rendering but remain difficult to obtain due to measurement complexity. This study investigates whether photogrammetry-reconstructed (PR) head and ear meshes, acquired with consumer hardware, can provide a practically useful baseline for individual HRTF synthesis. Using the SONICOM HRTF dataset, 72-image photogrammetry captures per subject were processed with Apple's Object Capture API to generate PR meshes for 150 subjects. Mesh2HRTF was used to compute PR synthetic HRTFs, which were compared against measured HRTFs, high-resolution 3D scan-derived HRTFs, KEMAR, and random HRTFs through numerical evaluation, auditory models, and a behavioural sound localisation experiment (N = 27). PR synthetic HRTFs preserved ITD cues but exhibited increased ILD and spectral errors. Auditory-model predictions and behavioural data showed substantially higher quadrant error rates, reduced elevation accuracy, and greater front-back confusions than measured HRTFs, performing worse than random HRTFs on perceptual metrics. Current photogrammetry pipelines support individual HRTF synthesis but are limited by insufficient pinna morphology details and high-frequency spectral fidelity needed for accurate individual HRTFs containing monaural cues.

空间音频3D重建声学建模摄影测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。