arXiv:2605.10398eess.AS2026-05中稿 · IWAENC 2026

用稀疏麦克风数据重建3D声场,提升空间音频建模精度。

SF-Flow: Sound field magnitude estimation via flow matching guided by sparse measurements

论文配图:SF-Flow: Sound field magnitude estimation via flow matching guided by sparse measurements
图 1 · 摘自论文原文
  • 基于流匹配的生成框架,结合3D U-Net与集合编码器
  • 在1kHz内实现高精度声场重构,训练速度更快
  • 适用于任意数量稀疏输入,适合声学空间建模场景

从稀疏麦克风测量中重建3D声场是一个基本但病态的问题,本文通过声学传输函数(ATF)幅度估计来解决。ATF幅度包含物理空间的关键听觉与声学特性,可用于房间表征与校正。尽管近期生成范式如流匹配(FM)在语音和音乐生成中表现优异,其在空间音频中的潜力仍待挖掘。我们提出一种新型3D ATF幅度重建框架,作为引导生成任务,采用由排列不变集合编码器条件化的3D U-Net。该架构可处理任意数量的稀疏输入,同时利用FM稳定高效的训练特性。实验表明,SF-Flow在高达1kHz的频率范围内实现高精度重建,训练速度显著快于自编码器基线,并随数据集规模增大持续提升性能。

原文摘要 · Abstract (English)

Reconstructing a 3D sound field from sparse microphone measurements is a fundamental yet ill-posed problem, which we address through Acoustic Transfer Function (ATF) magnitude estimation. ATF magnitude encapsulates key perceptual and acoustic properties of a physical space with applications in room characterization and correction. Although recent generative paradigms such as Flow Matching (FM) have achieved state-of-the-art performance in speech and music generation, their potential in spatial audio remains underexplored. We propose a novel framework for 3D ATF magnitude reconstruction as a guided generation task, with a 3D U-Net conditioned by a permutation-invariant set encoder. This architecture enables reconstruction from an arbitrary number of sparse inputs while leveraging the stable and efficient training properties of FM. Experimental results demonstrate that SF-Flow achieves accurate reconstruction up to \SI{1}{kHz}, trains substantially faster than the autoencoder baseline, and improves significantly with dataset size.

声场重建流匹配空间音频3D建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。