arXiv:2510.22439cs.SDcs.AI2025-10被引 1

用自然语言生成高保真混响响应,让虚拟声学环境更真实。

PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching

  • 先升频再生成:用变分自编码器提升分辨率,再用扩散模型根据描述生成混响。
  • 生成的混响误差仅8.8%,远低于基线的-37%。
  • 适合虚拟现实、建筑声学等需要灵活高质混响合成的场景。

房间冲激响应(RIR)生成对构建沉浸式虚拟声学环境至关重要。现有方法存在两大瓶颈:全频段RIR数据稀缺,且无法从多样化输入模态生成声学准确的响应。本文提出PromptReverb,一种两阶段生成框架:首先使用变分自编码器将带限RIR上采样至全频段(48 kHz),然后基于修正流匹配的条件扩散变换器模型,从自然语言描述生成RIR。实证评估表明,PromptReverb在感知质量和声学准确性上均优于现有方法,平均RT60误差为8.8%,而广泛使用的基线误差达-37%,生成参数更贴近真实房间特性。该方法可支持虚拟现实、建筑声学与音频制作中对灵活、高质量RIR合成的实际需求。

原文摘要 · Abstract (English)

Room impulse response (RIR) generation remains a critical challenge for creating immersive virtual acoustic environments. Current methods suffer from two fundamental limitations: the scarcity of full-band RIR datasets and the inability of existing models to generate acoustically accurate responses from diverse input modalities. We present PromptReverb, a two-stage generative framework that addresses these challenges. Our approach combines a variational autoencoder that upsamples band-limited RIRs to full-band quality (48 kHz), and a conditional diffusion transformer model based on rectified flow matching that generates RIRs from descriptions in natural language. Empirical evaluation demonstrates that PromptReverb produces RIRs with superior perceptual quality and acoustic accuracy compared to existing methods, achieving 8.8% mean RT60 error compared to -37% for widely used baselines and yielding more realistic room-acoustic parameters. Our method enables practical applications in virtual reality, architectural acoustics, and audio production where flexible, high-quality RIR synthesis is essential.

声学生成扩散模型多模态虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。