用大模型引导生成物理可信的动态图像,无需训练或标注。
PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM
- 用大模型评估生成动作的物理可信度,迭代优化运动提示。
- 在未训练情况下实现物理一致性,显著减少穿模和动量异常。
- 适合需要真实物理行为的3D内容生成场景,如动画与仿真。
近期3D内容生成技术推动了对兼具视觉真实性和物理一致性的动态模型的需求。然而,现有视频扩散模型常出现动量不守恒、物体穿模等不合理现象。现有物理感知方法多依赖特定任务微调或监督数据,难以扩展。为此,我们提出PhyMAGIC——一种无需训练的框架,仅凭单张图像即可生成物理一致的动态内容。该框架融合预训练图像到视频扩散模型、基于大模型的置信度引导推理以及可微分物理模拟器,直接生成可用于下游物理仿真的3D资产,无需微调或人工标注。通过利用大模型提供的置信度分数迭代优化运动提示,并结合模拟反馈进行修正,PhyMAGIC有效引导生成过程趋向物理合理动态。大量实验表明,PhyMAGIC优于当前主流视频生成器与物理感知基线,在提升物理属性推断与运动-文本对齐能力的同时,保持良好视觉保真度。
原文摘要 · Abstract (English)
Recent advances in 3D content generation have amplified demand for dynamic models that are both visually realistic and physically consistent. However, state-of-the-art video diffusion models frequently produce implausible results such as momentum violations and object interpenetrations. Existing physics-aware approaches often rely on task-specific fine-tuning or supervised data, which limits their scalability and applicability. To address the challenge, we present PhyMAGIC, a training-free framework that generates physically consistent motion from a single image. PhyMAGIC integrates a pre-trained image-to-video diffusion model, confidence-guided reasoning via LLMs, and a differentiable physics simulator to produce 3D assets ready for downstream physical simulation without fine-tuning or manual supervision. By iteratively refining motion prompts using LLM-derived confidence scores and leveraging simulation feedback, PhyMAGIC steers generation toward physically consistent dynamics. Comprehensive experiments demonstrate that PhyMAGIC outperforms state-of-the-art video generators and physics-aware baselines, enhancing physical property inference and motion-text alignment while maintaining visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。