用小波散射变换提升稀疏条件下声场重建精度
Masked Wavelet Scattering Transform Neural Field for Sound Field Reconstruction

- 用小波散射变换提取多尺度特征,引入统计先验
- 在HRTF上实现高保真上采样,优于基线方法
- 分两阶段学习掩码,适合个体化声场建模
本文提出一种基于小波散射变换(WST)的声场重建框架,利用WST作为多尺度特征提取器,在稀疏观测条件下引入统计先验。重建问题被建模为优化任务,通过神经场求解,并将WST嵌入训练损失函数。以头相关传递函数(HRTF)上采样为例进行验证,采用掩码策略对WST系数进行处理,形成两阶段流程:第一阶段从少量多主体数据中学习二值掩码;第二阶段将该掩码应用于个体HRTF的WST系数,以保留重建过程中的有效统计结构。与基线方法对比验证了该方法的有效性,同时作为组件消融分析,表明各模块贡献显著。
原文摘要 · Abstract (English)
In this paper, we propose a reconstruction framework that leverages the Wavelet Scattering Transform (WST) as a multi-scale feature extractor to impose statistical priors under sparse observation conditions. The reconstruction problem is formulated as an optimization task and solved using a neural field, with the WST incorporated into the training loss function. As a proof of concept, we validate the proposed method on HRTF upsampling. A masking strategy is applied to the WST coefficients, resulting in a two-phase procedure. The first phase learns a binary mask from a small multi-subject dataset, while the second phase applies the learned mask to the WST coefficients of an individual HRTF to preserve informative statistical structures during reconstruction. Validation against baseline methods, which also serve as an ablation study of the different components of the framework, demonstrates the effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。