arXiv:2511.20422cs.AIcs.CV2025-11被引 1

构建物理一致的声学-几何数据集,让模型理解物体形状如何发声。

VibraVerse: A Large-Scale Geometry-Acoustics Alignment Dataset for Physically-Consistent Multimodal Learning

  • 从3D几何出发,通过物理参数计算振动模态与声音信号
  • 在多个任务上实现比现有方法更高的准确率和可解释性
  • 适合研究物理感知、跨模态生成与具身智能的学者

理解物理世界需要基于物理规律的感知模型,而非单纯统计相关。现有跨模态学习框架多聚焦视觉与语言,缺乏物理一致性,忽视物体几何、材质、振动模式与声音之间的因果关系。本文提出VibraVerse,一个大规模几何-声学对齐数据集,明确建立从3D几何→物理属性→模态参数→声学信号的因果链。每个3D模型包含密度、杨氏模量、泊松比等显式物理属性及体素化几何,由此计算模态本征频率与本征向量,用于受控激励下的撞击声合成。为实现一致性对齐,提出CLASP对比学习框架,保留物理结构与声学响应间的因果对应。该框架确保样本间的一致性、可追溯性,并嵌入统一的形状、图像、声音表示空间。基于VibraVerse,定义了几何到声音预测、声音引导的形状重建、跨模态表征学习等基准任务。大量实验表明,训练于VibraVerse的模型在准确性、可解释性与泛化能力上表现更优。结果确立VibraVerse作为物理一致且因果可解释的跨模态学习基准,为声音引导的具身感知与物理世界深层理解提供基础。数据集将开源。

原文摘要 · Abstract (English)

Understanding the physical world requires perceptual models grounded in physical laws rather than mere statistical correlations. However, existing multimodal learning frameworks, focused on vision and language, lack physical consistency and overlook the intrinsic causal relationships among an object's geometry, material, vibration modes, and the sounds it produces. We introduce VibraVerse, a large-scale geometry-acoustics alignment dataset that explicitly bridges the causal chain from 3D geometry -> physical attributes -> modal parameters -> acoustic signals. Each 3D model has explicit physical properties (density, Young's modulus, Poisson's ratio) and volumetric geometry, from which modal eigenfrequencies and eigenvectors are computed for impact sound synthesis under controlled excitations. To establish this coherence, we introduce CLASP, a contrastive learning framework for cross-modal alignment that preserves the causal correspondence between an object's physical structure and its acoustic response. This framework enforces physically consistent alignment across modalities, ensuring that every sample is coherent, traceable to the governing equations, and embedded within a unified representation space spanning shape, image, and sound. Built upon VibraVerse, we define a suite of benchmark tasks for geometry-to-sound prediction, sound-guided shape reconstruction, and cross-modal representation learning. Extensive validations on these tasks demonstrate that models trained on VibraVerse exhibit superior accuracy, interpretability, and generalization across modalities. These results establish VibraVerse as a benchmark for physically consistent and causally interpretable multimodal learning, providing a foundation for sound-guided embodied perception and a deeper understanding of the physical world. The dataset will be open-sourced.

跨模态学习物理建模声学生成3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。