用少量数据生成逼真声学效果,提升虚拟环境沉浸感
Few-shot Acoustic Synthesis with Multimodal Flow Matching
- 基于流匹配的生成模型,从少量场景信息推断声学分布
- 仅需一次采样即超越八次采样的现有方法,性能提升显著
- 适用于低资源场景的声学建模,适合虚拟现实与空间音频应用
生成与场景一致的音频对沉浸式虚拟环境至关重要。现有神经声学场方法虽能实现空间连续的声音渲染,但依赖密集音频测量和高成本训练,仅适用于特定场景。少样本方法提升了跨场景扩展性,但仍需多段录音,且因确定性生成无法捕捉声学不确定性。本文提出流匹配声学生成(FLAC),一种概率性少样本声学合成方法,可在极简场景上下文条件下建模合理的房间冲击响应(RIR)分布。FLAC利用扩散变换器结合流匹配目标,在新场景任意位置生成RIR,条件为空间、几何与声学线索。在AcousticRooms和Hearing Anything Anywhere数据集上,其单次采样性能超越当前最佳的八次采样基线。为进一步评估生成质量,本文引入AGREE——一种联合声学-几何嵌入,通过检索与分布度量实现几何一致性评估。这是首次将生成式流匹配应用于显式RIR合成,为鲁棒、高效声学合成开辟新路径。
原文摘要 · Abstract (English)
Generating audio that is acoustically consistent with a scene is essential for immersive virtual environments. Recent neural acoustic field methods enable spatially continuous sound rendering but remain scene-specific, requiring dense audio measurements and costly training for each environment. Few-shot approaches improve scalability across rooms but still rely on multiple recordings and, being deterministic, fail to capture the inherent uncertainty of scene acoustics under sparse context. We introduce flow-matching acoustic generation (FLAC), a probabilistic method for few-shot acoustic synthesis that models the distribution of plausible room impulse responses (RIRs) given minimal scene context. FLAC leverages a diffusion transformer trained with a flow-matching objective to generate RIRs at arbitrary positions in novel scenes, conditioned on spatial, geometric, and acoustic cues. FLAC outperforms state-of-the-art eight-shot baselines with one-shot on both the AcousticRooms and Hearing Anything Anywhere datasets. To complement standard perceptual metrics, we further introduce AGREE, a joint acoustic-geometry embedding, enabling geometry-consistent evaluation of generated RIRs through retrieval and distributional metrics. This work is the first to apply generative flow matching to explicit RIR synthesis, establishing a new direction for robust and data-efficient acoustic synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。