提出可微分的多球散射模型,提升水下声源定位精度。
Analytical Exploration of Spatial Audio Cues: A Differentiable Multi-Sphere Scattering Model
- 基于半透明球体与刚性散射体构建解析式声波散射模型
- 在噪声环境下实现更优定位收敛,支持动态声源跟踪
- 适合做物理驱动的麦克风阵列系统,尤其水下场景
开发合成空间听觉系统的核心挑战在于精确建模声波散射。生物体通过身体对声波的散射产生依赖位置的双耳强度差(ILD)和时间差(ITD),从而实现三维空间听觉。传统头相关传输函数(HRTF)模型基于刚性散射,适用于陆地人类,但在水下失效,因水与软组织阻抗接近。受水生动物声学结构启发,我们提出一种新颖的、解析推导的闭合形式前向模型,用于模拟包含两个刚性球形散射体的半透明球体的声波散射。该模型能准确将声源方向、频率及材料属性映射为压力场,捕捉分层可穿透结构的复杂物理特性。关键优势在于模型完全可微分,可与机器学习算法结合优化目标函数以实现主动定位。我们验证了采用物理感知频率加权策略后,在噪声环境下的定位收敛性能提升,并通过扩展卡尔曼滤波器(EKF)实现了基于解析雅可比矩阵的精准移动声源追踪。研究表明,此类分层刚性与透明几何体的可微分散射模型,为利用散射线索而非传统波束成形的麦克风阵列提供了新基础,适用于陆地与水下场景。模型将开源发布。
原文摘要 · Abstract (English)
A primary challenge in developing synthetic spatial hearing systems, particularly underwater, is accurately modeling sound scattering. Biological organisms achieve 3D spatial hearing by exploiting sound scattering off their bodies to generate location-dependent interaural level and time differences (ITD/ILD). While Head-Related Transfer Function (HRTF) models based on rigid scattering suffice for terrestrial humans, they fail in underwater environments due to the near-impedance match between water and soft tissue. Motivated by the acoustic anatomy of underwater animals, we introduce a novel, analytically derived, closed-form forward model for scattering from a semi-transparent sphere containing two rigid spherical scatterers. This model accurately maps source direction, frequency, and material properties to the pressure field, capturing the complex physics of layered, penetrable structures. Critically, our model is implemented in a fully differentiable setting, enabling its integration with a machine learning algorithm to optimize a cost function for active localization. We demonstrate enhanced convergence for localization under noise using a physics-informed frequency weighting scheme, and present accurate moving-source tracking via an Extended Kalman Filter (EKF) with analytically computed Jacobians. Our work suggests that differentiable models of scattering from layered rigid and transparent geometries offer a promising new foundation for microphone arrays that leverage scattering-based spatial cues over conventional beamforming, applicable to both terrestrial and underwater applications. Our model will be made open source.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。