arXiv:2608.05145cs.CVcs.MM2026-08

用少量敲击声和多视角图像重建物体的声学特性,实现高保真音效生成。

Objects as Audio-Visual Modal Sound Fields

论文配图:Objects as Audio-Visual Modal Sound Fields
图 1 · 摘自论文原文
  • 结合3D高斯点云与视觉特征,用少数音频样本建模物体声学响应。
  • 在两个真实数据集上音效渲染精度优于物理模拟与纯数据方法。
  • 支持接触位置定位与声音编辑,适合具身智能与交互系统应用。

现代3D重建擅长建模物体几何与外观,却忽略物理交互带来的丰富声学线索。物体敲击声传递材质、刚度与结构信息,可补充视觉感知。现有方法或依赖昂贵物理仿真,或需大规模数据训练。本文提出音频-视觉模态声场(AV-MSF),仅需少量敲击录音与多视角图像,即可在物体层面重建声学表征。AV-MSF基于3D高斯点云融合密集视觉特征,提供强几何先验,并以紧凑且具物理意义的模态参数表示冲击声场,实现鲁棒的少样本重建。在两个真实世界数据集上的实验表明,该方法在冲击声渲染上达到当前最优性能,超越物理驱动与数据驱动基线。此外,我们展示了该表征支持的下游应用,包括接触定位与声音编辑。

原文摘要 · Abstract (English)

While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.

3D重建音频建模少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。