arXiv:2509.05205eess.AScs.SD2025-09中稿 · ASRU 2025被引 5

融合音视频文本信息,提升混响声响应估计精度

MEAN-RIR: Multi-Modal Environment-Aware Network for Robust Room Impulse Response Estimation

  • 多模态输入通过交叉注意力融合环境信息
  • 分离生成直达声/早期反射与晚期混响,提升重建质量
  • 适合需要精准声学建模的语音增强与虚拟现实场景

本文提出一种多模态环境感知网络MEAN-RIR,基于音频、视觉和文本多源环境信息预测房间冲激响应(RIR)。以混响语音作为主输入,结合全景图像和文本描述作为辅助输入。各模态分别经编码器处理后,通过交叉注意力模块实现跨模态交互。解码器生成两部分:第一部分捕捉直达声与早期反射,第二部分生成掩码以调制可学习滤波噪声,合成晚期混响。二者混合重建最终RIR。实验表明,该方法显著提升RIR估计性能,尤其在声学参数上表现优异。

原文摘要 · Abstract (English)

This paper presents a Multi-Modal Environment-Aware Network (MEAN-RIR), which uses an encoder-decoder framework to predict room impulse response (RIR) based on multi-level environmental information from audio, visual, and textual sources. Specifically, reverberant speech capturing room acoustic properties serves as the primary input, which is combined with panoramic images and text descriptions as supplementary inputs. Each input is processed by its respective encoder, and the outputs are fed into cross-attention modules to enable effective interaction between different modalities. The MEAN-RIR decoder generates two distinct components: the first component captures the direct sound and early reflections, while the second produces masks that modulate learnable filtered noise to synthesize the late reverberation. These two components are mixed to reconstruct the final RIR. The results show that MEAN-RIR significantly improves RIR estimation, with notable gains in acoustic parameters.

声学建模多模态语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。