arXiv:2606.30114eess.AS2026-06中稿 · and presented at t…被引 1

对比五种头相关传输函数,发现合成方法影响听觉定位精度。

Evaluation of Head-Related Transfer Functions Across Five Levels of Individualisation in Virtual Reality

论文配图:Evaluation of Head-Related Transfer Functions Across Five Levels of Individualisation in Virtual Reality
图 1 · 摘自论文原文
  • 在虚拟现实里测试五类HRTF:实测、通用、随机、高分辨率合成和摄影测量合成。
  • 合成方法中,高分辨率模型表现接近真实测量,摄影测量法最差。
  • 随机选择的HRTF反而比通用模型(KEMAR)更准,适合快速部署。

头相关传输函数(HRTFs)是虚拟与增强现实系统中空间听觉的基础。尽管个体化HRTF能捕捉听众特异性形态,但其实际限制导致广泛使用通用HRTF及合成方法兴起。然而,这些方法的感知影响尚未在单一研究中系统比较。本研究分析了19名受试者完成的两次虚拟现实声音定位实验数据,采用交错设计,对比五种条件:实测个体化、KEMAR通用、随机非个体化、高分辨率扫描合成与摄影测量合成的HRTF。跨会话测试-重测稳定性支持将差异归因于感知而非实验效应。结果表明,横向定位指标对HRTF类型不敏感;而极坐标域指标与混淆率则高度依赖于HRTF类型。随机选取的HRTF在多项极坐标指标上优于KEMAR。高分辨率合成的性能接近实测个体化水平,而摄影测量合成与KEMAR均表现出最大退化。研究为非个体化基线选择提供依据,并强调数值合成中网格分辨率对仰角定位任务的重要性。

原文摘要 · Abstract (English)

Head-related transfer functions (HRTFs) underpin spatial hearing in virtual and augmented reality systems. Whilst individual HRTFs capture listener-specific morphology, their practical limitations have led to widespread use of generic HRTFs and growing interest in synthetic approaches. Yet their relative perceptual impact remains rarely compared within a single study. In this study, we analysed data from 19 listeners that completed two virtual reality sound localisation experiments with complementary subsets of interleaved HRTF conditions enabling within-subject comparison of five conditions: individually measured, KEMAR, randomly selected non-individual measured, high-resolution scan-based synthetic and photogrammetry-based synthetic HRTFs. Test-retest stability of the individually measured baseline across sessions supported pooling across experiments and attributing differences to perceptual rather than session effects. Across HRTF conditions, lateral localisation metrics were largely insensitive to HRTF type, whereas polar-domain metrics and confusion rates showed strong HRTF dependence. Random HRTFs outperformed KEMAR on several polar metrics. High-resolution synthetic HRTFs matched individual measured performance, whilst photogrammetry-based synthetic HRTFs, alongside KEMAR, showed the greatest degradation. These findings clarify practical choices for non-individual baselines and highlight the importance of mesh resolution when using numerical synthesis for elevation-dependent localisation tasks.

虚拟现实声学定位合成模型听觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。