arXiv:2605.00431cs.SDcs.CV2026-05中稿 · CVPR

利用视觉音频模型实现去混响与房间响应估计,无需修改网络结构。

MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation

论文配图:MMAudioReverbs: Video-Guided Acoustic Modeling for Dereverberation and Room Impulse Response Estimation
图 1 · 摘自论文原文
  • 基于预训练视觉音频模型,无需改动架构即可处理声学问题。
  • 在小数据集微调后,视听线索在不同声学条件下各有优势。
  • 为大模型在物理声学分析中的应用提供了新思路,适合音频处理研究者。

尽管近期的视觉到音频(V2A)模型在从视觉输入生成语义合理的声音方面表现优异,但它们并未显式建模混响或房间脉冲响应(RIR)等声学效应,因此对这些效应的控制能力有限。然而,我们假设这些V2A模型隐式地掌握了空间音频与相应视觉线索之间的关系。本文重新审视了一个V2A模型,提出利用预训练模型作为物理基础声学处理的先验知识。基于当前最先进的V2A模型MMAudio,我们提出了MMAudioReverbs,一个统一框架,可同时处理i) 去混响和ii) 房间脉冲响应(RIR)估计,且无需网络结构修改,并在小规模数据集上进行微调。实验结果表明,音频与视觉线索在不同类型的物理声学条件下分别具有优势,这表明基础V2A模型可用于物理驱动的声学分析。

原文摘要 · Abstract (English)

Although recent video-to-audio (V2A) models excelled at synthesizing semantically plausible sounds from visual inputs, they do not explicitly model room-acoustic effects such as reverberation or room impulse responses (RIRs), and thus offer limited controllability over these effects. However, we hypothesize that such V2A models implicitly have semantic knowledge of the relationship between spatial audio and the corresponding vision cues. In this paper, we revisit a V2A model for the sake of the above, and propose the way to utilize the pretrained model as prior for physically grounded room-acoustic processing. Based on one of the state-of-the-art V2A models, MMAudio, we propose MMAudioReverbs that is a unified framework dealing with i) dereverberation and ii) room impulse response (RIR) estimation without network architectural modification, and fine-tuned on a small dataset. Experimental results showed that audio and visual cues respectively have advantage depending on the type of physical room acoustics. It implies that foundation V2A models can be used for physically grounded room-acoustic analysis.

音频处理视觉音频去混响声学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。