根据材料配置生成室内声音效果,让虚拟场景听感更真实。
How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
- 用音视频信息编码场景,动态生成目标混响响应
- 在新数据集上生成的混响响应质量优于多个基准方法
- 适合做虚拟现实、声学设计等需要精准听感模拟的场景
如果一个录音室地板铺地毯、墙面贴吸音板,声音会如何变化?我们提出材料可控的声学特征生成任务:给定一个具有特定音视频特性的室内场景,用户可在推理时指定材料配置,模型生成对应的房间冲激响应(RIR)。为此,我们设计了一种新型编码器-解码器结构,从音视频观测中提取场景关键属性,并根据用户提供的材料参数生成目标RIR。该模型能基于动态定义的材料组合生成多样化高保真RIR。为支持此任务,我们构建了新的基准数据集Acoustic Wonderland Dataset,用于在多样且复杂的条件下开发和评估材料感知的RIR预测方法。实验结果表明,所提模型能有效编码材料信息,生成的RIR在多个指标上超越基线与先进方法。
原文摘要 · Abstract (English)
How would the sound in a studio change with a carpeted floor and acoustic tiles on the walls? We introduce the task of material-controlled acoustic profile generation, where, given an indoor scene with specific audio-visual characteristics, the goal is to generate a target acoustic profile based on a user-defined material configuration at inference time. We address this task with a novel encoder-decoder approach that encodes the scene's key properties from an audio-visual observation and generates the target Room Impulse Response (RIR) conditioned on the material specifications provided by the user. Our model enables the generation of diverse RIRs based on various material configurations defined dynamically at inference time. To support this task, we create a new benchmark, the Acoustic Wonderland Dataset, designed for developing and evaluating material-aware RIR prediction methods under diverse and challenging settings. Our results demonstrate that the proposed model effectively encodes material information and generates high-fidelity RIRs, outperforming several baselines and state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。