arXiv:2509.02571eess.AScs.AI2025-09

用物理感知核函数提升声场方向向量的少样本建模精度。

Gaussian Process Regression of Steering Vectors With Physics-Aware Deep Composite Kernels for Augmented Listening

  • 结合高斯过程与神经场,构建物理感知复合核函数。
  • 仅需不到十倍测量数据即达最优性能,抗过拟合能力更强。
  • 适合声学系统设计、空间音频渲染等需要少样本建模的场景。

本文研究了频率和麦克风/声源位置上连续的波束成形向量表示,用于增强听觉(如空间滤波和双耳渲染),支持用户参数化控制重现声场。传统波束成形向量在理想环境中采用代数表示,无法处理声场散射效应。通常需在专用设施中采集离散实测向量并进行超分辨率重建。近期物理感知深度学习方法有效解决了此问题,但确定性超分辨率在测量空间非均匀不确定性下易过拟合。为此,本文将神经场(NF)表达与高斯过程(GP)的严谨概率框架结合,提出一种物理感知复合核函数,建模入射波方向及后续散射效应。综合对比实验表明,该方法在数据不足条件下仍具有效性。在模拟SPEAR挑战数据的语音增强与双耳渲染任务中,仅需不足十倍的测量次数即可达到基准性能。

原文摘要 · Abstract (English)

This paper investigates continuous representations of steering vectors over frequency and microphone/source positions for augmented listening (e.g., spatial filtering and binaural rendering), enabling user-parameterized control of the reproduced sound field. Steering vectors have typically been used for representing the spatial response of a microphone array as a function of the look-up direction. The basic algebraic representation of these quantities assuming an idealized environment cannot deal with the scattering effect of the sound field. One may thus collect a discrete set of real steering vectors measured in dedicated facilities and super-resolve (i.e., upsample) them. Recently, physics-aware deep learning methods have been effectively used for this purpose. Such deterministic super-resolution, however, suffers from the overfitting problem due to the non-uniform uncertainty over the measurement space. To solve this problem, we integrate an expressive representation based on the neural field (NF) into the principled probabilistic framework based on the Gaussian process (GP). Specifically, we propose a physics-aware composite kernel that models the directional incoming waves and the subsequent scattering effect. Our comprehensive comparative experiment showed the effectiveness of the proposed method under data insufficiency conditions. In downstream tasks such as speech enhancement and binaural rendering using the simulated data of the SPEAR challenge, the oracle performances were attained with less than ten times fewer measurements.

声场建模高斯过程物理感知少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。