单图生成物理真实运动视频,融合力学模拟与扩散模型。
PhysMotion: Physics-Grounded Dynamics From a Single Image
- 用连续介质力学模拟构建3D高斯表示的初始动力学
- 通过可微MPM实现弹性塑性材料的物理演算,保持运动一致性
- 结合文本到图像扩散模型提升细节与时空连贯性,适合影视动画生成
我们提出PhysMotion,一种新颖框架,利用基于原理的物理仿真引导从单张图像和输入条件(如受力与力矩)生成的中间3D表示,从而生成高质量、物理合理的视频。该方法通过连续介质力学为基础的模拟作为先验知识,克服传统数据驱动生成模型在物理一致性上的局限。首先,通过几何优化从单图重建前向3D高斯表示;随后,使用可微分的材料点法(MPM)结合连续介质力学中的弹塑性模型进行时间步进,提供真实动态的基础,尽管细节较粗略。为增强几何、外观并确保时空一致性,我们采用带跨帧注意力的文本到图像(T2I)扩散模型对初始模拟结果进行精细化处理,最终生成保留输入图像复杂细节且物理合理的视频。我们进行了全面的定性和定量评估以验证方法有效性。
原文摘要 · Abstract (English)
We introduce PhysMotion, a novel framework that leverages principled physics-based simulations to guide intermediate 3D representations generated from a single image and input conditions (e.g., applied force and torque), producing high-quality, physically plausible video generation. By utilizing continuum mechanics-based simulations as a prior knowledge, our approach addresses the limitations of traditional data-driven generative models and result in more consistent physically plausible motions. Our framework begins by reconstructing a feed-forward 3D Gaussian from a single image through geometry optimization. This representation is then time-stepped using a differentiable Material Point Method (MPM) with continuum mechanics-based elastoplasticity models, which provides a strong foundation for realistic dynamics, albeit at a coarse level of detail. To enhance the geometry, appearance and ensure spatiotemporal consistency, we refine the initial simulation using a text-to-image (T2I) diffusion model with cross-frame attention, resulting in a physically plausible video that retains intricate details comparable to the input image. We conduct comprehensive qualitative and quantitative evaluations to validate the efficacy of our method. Our project page is available at: https://supertan0204.github.io/physmotion_website/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。