arXiv:2511.20562cs.CV2025-11被引 8

让视频生成具备物理可控性,从单图生成真实动态行为

PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding

  • 分部件重建物体初始物理属性,实现精准控制
  • 通过时序指令与可编辑模拟,生成高保真动态视频
  • 适合需要物理真实性的视频生成场景

当前视频生成模型虽有较高视觉质量,但缺乏显式的物理可控性与合理性。尽管已有研究尝试用基于物理的渲染引导生成,但仍面临复杂物理属性建模困难及长时间序列行为控制不佳的问题。本文提出PhysChoreo框架,仅需一张图像即可生成具有多样可控性与物理真实感的视频。方法分为两阶段:首先通过分部件感知的物理属性重建,估计图像中所有物体的静态初始物理特性;随后在时序指令与可编辑物理模拟驱动下,合成高质量、具丰富动态行为的视频。实验表明,PhysChoreo在多个评估指标上优于现有最先进方法,能有效生成兼具物理真实性和复杂动态行为的视频。

原文摘要 · Abstract (English)

While recent video generation models have achieved significant visual fidelity, they often suffer from the lack of explicit physical controllability and plausibility. To address this, some recent studies attempted to guide the video generation with physics-based rendering. However, these methods face inherent challenges in accurately modeling complex physical properties and effectively control ling the resulting physical behavior over extended temporal sequences. In this work, we introduce PhysChoreo, a novel framework that can generate videos with diverse controllability and physical realism from a single image. Our method consists of two stages: first, it estimates the static initial physical properties of all objects in the image through part-aware physical property reconstruction. Then, through temporally instructed and physically editable simulation, it synthesizes high-quality videos with rich dynamic behaviors and physical realism. Experimental results show that PhysChoreo can generate videos with rich behaviors and physical realism, outperforming state-of-the-art methods on multiple evaluation metrics.

视频生成物理模拟可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。