arXiv:2603.12442eess.AS2026-03

用扩散模型补全声学脉冲响应,更真实且无需固定输入时长

Room Impulse Response Completion Using Signal-Prediction Diffusion Models Conditioned on Simulated Early Reflections

  • 以几何模拟的早期反射为条件,用扩散模型补全完整脉冲响应
  • 在早期响应补全和能量衰减曲线重建上优于现有最佳方法
  • 适合做沉浸式音频、声学增强等需要真实脉冲响应的研究者

房间脉冲响应(RIR)是音频数据增强、声学信号处理和沉浸式音频渲染的基础。虽然几何模拟器(如图像源法,ISM)能高效生成早期反射,但因缺少声波效应而缺乏实测RIR的真实感。本文提出一种基于扩散模型的RIR补全方法,以ISM模拟的直达声与早期反射为条件进行信号预测。不同于现有方法,该方法对输入早期反射无固定时长限制。进一步引入无分类器引导,使生成结果向使用Treble SDK物理仿真得到的真实RIR分布靠拢。客观评估表明,该方法在早期RIR补全和能量衰减曲线重建方面优于当前最优基线。

原文摘要 · Abstract (English)

Room impulse responses (RIRs) are fundamental to audio data augmentation, acoustic signal processing, and immersive audio rendering. While geometric simulators such as the image source method (ISM) can efficiently generate early reflections, they lack the realism of measured RIRs due to missing acoustic wave effects. We propose a diffusion-based RIR completion method using signal-prediction conditioned on ISM-simulated direct-path and early reflections. Unlike state-of-the-art methods, our approach imposes no fixed duration constraint on the input early reflections. We further incorporate classifier-free guidance to steer generation toward a target distribution learned from physically realistic RIRs simulated with the Treble SDK. Objective evaluation demonstrates that the proposed method outperforms a state-of-the-art baseline in early RIR completion and energy decay curve reconstruction.

脉冲响应扩散模型音频生成声学模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。