arXiv:2604.08967cs.SD2026-04

用频谱图构建声场,实现高保真空间音频重建

AudioGS: Spectrogram-Based Audio Gaussian Splatting for Sound Field Reconstruction

  • 将声场建模为带方向与衰减特性的音频高斯点
  • 在Replay-NVAS数据集上降低14%幅度误差,减少25%感知质量误差
  • 无需视觉信息,适合无摄像头环境下的空间音频生成

空间音频是沉浸式虚拟体验的核心,但从稀疏观测中合成高保真双耳音频仍具挑战。现有方法多依赖视觉先验的隐式神经表示,难以捕捉精细声学结构。受3D高斯溅射(3DGS)启发,我们提出AudioGS,一种无需视觉信息的新型框架,通过频谱图显式编码声场为一组音频高斯点。每个时频帧对应一个音频高斯,配备双球谐函数(SH)系数和衰减系数。针对目标听觉位置,通过评估SH场捕捉方向性,结合几何引导的距离衰减与相位校正,重建波形。在Replay-NVAS数据集上的实验表明,AudioGS成功捕捉复杂空间线索,优于最先进的视觉依赖基线。具体而言,相比表现最佳的视觉引导方法,音频幅度重建误差(MAG)降低超过14%,感知质量指标(DPAM)降低约25%。

原文摘要 · Abstract (English)

Spatial audio is fundamental to immersive virtual experiences, yet synthesizing high-fidelity binaural audio from sparse observations remains a significant challenge. Existing methods typically rely on implicit neural representations conditioned on visual priors, which often struggle to capture fine-grained acoustic structures. Inspired by 3D Gaussian Splatting (3DGS), we introduce AudioGS, a novel visual-free framework that explicitly encodes the sound field as a set of Audio Gaussians based on spectrograms. AudioGS associates each time-frequency bin with an Audio Gaussian equipped with dual Spherical Harmonic (SH) coefficients and a decay coefficient. For a target pose, we render binaural audio by evaluating the SH field to capture directionality, incorporating geometry-guided distance attenuation and phase correction, and reconstructing the waveform. Experiments on the Replay-NVAS dataset demonstrate that AudioGS successfully captures complex spatial cues and outperforms state-of-the-art visual-dependent baselines. Specifically, AudioGS reduces the magnitude reconstruction error (MAG) by over 14% and reduces the perceptual quality metric (DPAM) by approximately 25% compared to the best performing visual-guided method.

空间音频声场重建高斯溅射频谱建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。