arXiv:2510.02813eess.AS2025-10中稿 · poster presentatio…被引 1

用图神经网络提升低分辨率照片建模的耳廓精度,实现高保真声音定位重建

Enhancing Photogrammetry Reconstruction For HRTF Synthesis Via A Graph Neural Network

  • 通过图神经网络结合神经细分,将低精度照片重建头像升维
  • 生成的高精度网格可合成接近真实测量值的个体化HRTF
  • 适合需要低成本个性化音频系统的研究者与开发者

传统头相关传递函数(HRTF)获取依赖专业设备和声学知识,难以普及。高分辨率3D建模虽可数值合成HRTF,但高端3D扫描仪成本高、难获取。摄影测量法可生成3D头像网格,但分辨率不足限制其用于HRTF合成。本研究探索使用图神经网络(GNN)与神经细分技术,将低分辨率摄影测量重建(PR)网格上采样为高分辨率网格,以支持个体化HRTF合成。基于SONICOM数据集,利用Apple Photogrammetry API生成低分辨率头像网格,构建配对的高低分辨率网格数据集,训练GNN模型,采用基于豪斯多夫距离的损失函数。在未见过的摄影测量数据上验证模型几何精度及通过Mesh2HRTF合成的HRTF性能。合成的HRTF与高分辨率扫描计算结果、实测HRTF以及KEMAR HRTF进行对比,采用感知相关数值分析和行为实验(包括定位任务与空间掩蔽释放,SRM)评估。

原文摘要 · Abstract (English)

Traditional Head-Related Transfer Functions (HRTFs) acquisition methods rely on specialised equipment and acoustic expertise, posing accessibility challenges. Alternatively, high-resolution 3D modelling offers a pathway to numerically synthesise HRTFs using Boundary Elements Methods and others. However, the high cost and limited availability of advanced 3D scanners restrict their applicability. Photogrammetry has been proposed as a solution for generating 3D head meshes, though its resolution limitations restrict its application for HRTF synthesis. To address these limitations, this study investigates the feasibility of using Graph Neural Networks (GNN) using neural subdivision techniques for upsampling low-resolution Photogrammetry-Reconstructed (PR) meshes into high-resolution meshes, which can then be employed to synthesise individual HRTFs. Photogrammetry data from the SONICOM dataset are processed using Apple Photogrammetry API to reconstruct low-resolution head meshes. The dataset of paired low- and high-resolution meshes is then used to train a GNN to upscale low-resolution inputs to high-resolution outputs, using a Hausdorff Distance-based loss function. The GNN's performance on unseen photogrammetry data is validated geometrically and through synthesised HRTFs generated via Mesh2HRTF. Synthesised HRTFs are evaluated against those computed from high-resolution 3D scans, to acoustically measured HRTFs, and to the KEMAR HRTF using perceptually-relevant numerical analyses as well as behavioural experiments, including localisation and Spatial Release from Masking (SRM) tasks.

三维重建图神经网络音频合成声音定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。