arXiv:2506.06440cs.GRcs.CV2025-06CVPR被引 20

用视频重建物体外观、形状和物理属性,高效且通用。

Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation

  • 基于线性骨骼绑定的无网格简化模拟,快速估计物理状态。
  • 仅需几分钟优化即可精准匹配视频观测,精度优于传统方法。
  • 适合需要快速高保真物理重建的仿真与动画应用。

从视频中准确重建带纹理的物体形状与物理属性是一项极具挑战的任务。现有方法多依赖可微分模拟器与渲染器的复杂优化流程,需为每场景调参且计算成本高,限制了实用性与泛化能力。本文提出Vid2Sim,一种通用性强的视频驱动重建框架,采用基于线性骨骼绑定(LBS)的无网格简化模拟,兼具高效计算与灵活表征能力。首先通过前馈神经网络从视频中重建物理系统状态,再以轻量级优化流程在数分钟内精细调整外观、几何与物理属性,使其与视频高度一致。重建后,系统可实现高质量、高效率的无网格仿真。大量实验表明,该方法在几何与物理属性重建上兼具更高精度与效率。

原文摘要 · Abstract (English)

Faithfully reconstructing textured shapes and physical properties from videos presents an intriguing yet challenging problem. Significant efforts have been dedicated to advancing such a system identification problem in this area. Previous methods often rely on heavy optimization pipelines with a differentiable simulator and renderer to estimate physical parameters. However, these approaches frequently necessitate extensive hyperparameter tuning for each scene and involve a costly optimization process, which limits both their practicality and generalizability. In this work, we propose a novel framework, Vid2Sim, a generalizable video-based approach for recovering geometry and physical properties through a mesh-free reduced simulation based on Linear Blend Skinning (LBS), offering high computational efficiency and versatile representation capability. Specifically, Vid2Sim first reconstructs the observed configuration of the physical system from video using a feed-forward neural network trained to capture physical world knowledge. A lightweight optimization pipeline then refines the estimated appearance, geometry, and physical properties to closely align with video observations within just a few minutes. Additionally, after the reconstruction, Vid2Sim enables high-quality, mesh-free simulation with high efficiency. Extensive experiments demonstrate that our method achieves superior accuracy and efficiency in reconstructing geometry and physical properties from video data.

视频重建物理仿真无网格模拟神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。