用多模态模型和视频扩散实现更真实的4D物理场景动态模拟
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
- 通过图像查询自动识别材料类型并初始化参数
- 结合视频扩散与可微分MPM实现高精度动态交互模拟
- 适合需要真实物理演化的影视/游戏/工程仿真场景
真实动态场景的模拟需准确捕捉多种材料属性并建模复杂的物体相互作用,但现有方法仅限于基础材料类型且可调参数有限,难以表达现实材料的复杂性。我们提出PhysFlow,利用多模态基础模型和视频扩散,实现增强的4D动态场景模拟。该方法通过图像查询识别材料类型并初始化材料参数,同时推断3D Gaussian splats以实现精细场景表示。进一步采用带可微分材料点法(MPM)和光流引导的视频扩散优化材料参数,而非依赖渲染损失或得分蒸馏采样(SDS)损失。该集成框架可准确预测并真实模拟现实场景中的动态交互,提升了基于物理模拟的精度与灵活性。
原文摘要 · Abstract (English)
Realistic simulation of dynamic scenes requires accurately capturing diverse material properties and modeling complex object interactions grounded in physical principles. However, existing methods are constrained to basic material types with limited predictable parameters, making them insufficient to represent the complexity of real-world materials. We introduce PhysFlow, a novel approach that leverages multi-modal foundation models and video diffusion to achieve enhanced 4D dynamic scene simulation. Our method utilizes multi-modal models to identify material types and initialize material parameters through image queries, while simultaneously inferring 3D Gaussian splats for detailed scene representation. We further refine these material parameters using video diffusion with a differentiable Material Point Method (MPM) and optical flow guidance rather than render loss or Score Distillation Sampling (SDS) loss. This integrated framework enables accurate prediction and realistic simulation of dynamic interactions in real-world scenarios, advancing both accuracy and flexibility in physics-based simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。