arXiv:2507.12136cs.SDeess.AS2025-07中稿 · IEEE Workshop on A…被引 6

用声音特性生成房间混响,让虚拟空间更真实

Room Impulse Response Generation Conditioned on Acoustic Parameters

  • 直接以混响时间等听觉参数为条件生成混响响应
  • MaskGIT模型在客观与主观测试中表现最优
  • 适合需要听感真实而非几何精确的音频应用

基于深度神经网络生成房间混响响应(RIR)在虚拟现实、音频后期制作等领域日益受到关注。现有方法多依赖房间尺寸、形状和表面材料等几何信息,但在布局未知或听觉真实感更重要的场景下受限。本文提出新策略:直接以一组RIR声学参数为条件,包括宽带与频带混响时间、直达声与混响比等。通过指定空间的听觉特征而非几何形态,实现更灵活、感知驱动的RIR生成。研究对比了四种模型:自回归Transformer、MaskGIT、流匹配模型及基于分类器的方法,均在描述音频编码域中运行。客观与主观评估表明,所提方法性能达到或超过当前最优水平,其中MaskGIT表现最佳。

原文摘要 · Abstract (English)

The generation of room impulse responses (RIRs) using deep neural networks has attracted growing research interest due to its applications in virtual and augmented reality, audio postproduction, and related fields. Most existing approaches condition generative models on physical descriptions of a room, such as its size, shape, and surface materials. However, this reliance on geometric information limits their usability in scenarios where the room layout is unknown or when perceptual realism (how a space sounds to a listener) is more important than strict physical accuracy. In this study, we propose an alternative strategy: conditioning RIR generation directly on a set of RIR acoustic parameters. These parameters include various measures of reverberation time and direct sound to reverberation ratio, both broadband and bandwise. By specifying how the space should sound instead of how it should look, our method enables more flexible and perceptually driven RIR generation. We explore both autoregressive and non-autoregressive generative models operating in the Descript Audio Codec domain, using either discrete token sequences or continuous embeddings. Specifically, we have selected four models to evaluate: an autoregressive transformer, the MaskGIT model, a flow matching model, and a classifier-based approach. Objective and subjective evaluations are performed to compare these methods with state-of-the-art alternatives. Results show that the proposed models match or outperform state-of-the-art alternatives, with the MaskGIT model achieving the best performance.

声学建模生成模型音频合成感知真实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。