arXiv:2603.02508eess.AScs.SD2026-03

拆解声学建模对个人声区渲染的影响,指导有限预算下训练数据构建。

Decomposing the Influence of Physical Acoustic Modeling on Neural Personal Sound Zone Rendering: An Ablation Study

  • 逐项加入扬声器频响、指向性与头相关传递函数,系统评估各成分贡献。
  • 指向性模型提升声区分离10.05 dB,是改善隔离效果的关键因素。
  • 适用于音频系统设计与神经渲染研究者,尤其关注真实场景泛化性能。

基于深度学习的个人声区(PSZ)系统依赖模拟的声学传递函数(ATFs)进行训练,但理想化的点源模型存在显著的仿真到现实差距。尽管引入物理先验可提升泛化能力,但各组件的具体贡献仍不明确。本文针对一种头姿态相关的双耳PSZ渲染器(BSANN),开展受控消融实验,逐步在模拟ATFs中加入三类物理成分:(i) 实测扬声器频响(FR)、(ii) 解析圆柱活塞指向性(DIR)、(iii) 刚性球头相关传递函数(RS-HRTF)。通过两个假人头的现场测量,评估四种配置在100–20000 Hz范围内的性能,指标包括区间隔离度(IZI)、程序间干扰(IPI)和串扰抑制(XTC)。结果表明:FR实现频谱校准,带来适度的XTC改善和听众间IPI不平衡降低;DIR提供最稳定的声区分离增益(平均IZI/IPI提升10.05 dB);RS-HRTF主导双耳分离效果,使XTC提升+2.38/+2.89 dB(平均从4.51提升至7.91 dB),主要体现在2 kHz以上频段,但引入轻微听者依赖的IZI/IPI波动。研究为资源受限时训练ATFs的建模优先级提供了明确指导。

原文摘要 · Abstract (English)

Deep learning-based Personal Sound Zones (PSZs) rely on simulated acoustic transfer functions (ATFs) for training, yet idealized point-source models exhibit large sim-to-real gaps. While physically informed components improve generalization, individual contributions remain unclear. This paper presents a controlled ablation study on a head-pose-conditioned binaural PSZ renderer using the Binaural Spatial Audio Neural Network (BSANN). We progressively enrich simulated ATFs with three components: (i) anechoically measured frequency responses of the particular loudspeakers(FR), (ii) analytic circular-piston directivity (DIR), and (iii) rigid-sphere head-related transfer functions (RS-HRTF). Four configurations are evaluated via in-situ measurements with two dummy heads. Performance metrics include inter-zone isolation (IZI), inter-program interference (IPI), and crosstalk cancellation (XTC) over 100-20000 Hz. Results show FR provides spectral calibration, yielding modest XTC improvements and reduced inter-listener IPI imbalance. DIR delivers the most consistent sound-zone separation gains (10.05 dB average IZI/IPI). RS-HRTF dominates binaural separation, boosting XTC by +2.38/+2.89 dB (average 4.51 to 7.91 dB), primarily above 2 kHz, while introducing mild listener-dependent IZI/IPI shifts. These findings guide prioritization of measurements and models when constructing training ATFs under limited budgets.

声学建模神经渲染个人声区

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。