让视频生成更懂物理:通过反复试错优化动作控制
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation

- 用可执行的物理程序作为假设,逐步验证修正
- 在复杂轨迹和多阶段交互上提升物理真实性
- 适合需要精确物理控制的动画与仿真场景
近期基于物理的视频生成利用物理模拟作为先验,引导视频合成向物理合理结果靠拢。模拟过程由物理参数控制,这些参数通常由视觉语言模型一次性生成。然而,单次预测常难以准确将用户意图转化为可执行的模拟,尤其在细粒度物体动态、复杂运动轨迹和时序结构化交互方面表现不佳。本文提出PhysAgent,一种反射式智能体框架,实现了物理程序生成、物理模拟、阶段验证与针对性修复之间的闭环。该框架不仅提升了耦合物理参数的控制能力,还通过将每个物理程序视为可执行假设,逐步实现复杂轨迹、多阶段交互与精确事件结果。此外,我们设计了一套物理控制API,支持更稳定复杂的运动行为。大量实验表明,PhysAgent生成的视频更具物理合理性,提示对齐效果更好,并在多样化物理场景中具备更强泛化能力。
原文摘要 · Abstract (English)
Recent advances in physics-grounded video generation leverage physics simulation as a physical prior to guide video synthesis toward physically plausible outcomes. The simulation process is controlled by physical specifications, which are typically generated by a vision-language model in a single pass. Such one-shot prediction often fails to accurately translate user intent into executable simulations, particularly for fine-grained object dynamics, complex motion trajectories, and temporally structured interactions. In this paper, we propose PhysAgent, a reflective agentic framework that closes the loop among physical program generation, physics simulation, stage-specific verification, and targeted program repair. Beyond improving the control of coupled physical parameters, our framework enables the agent to progressively realize complex trajectories, multi-stage interactions, and precise event outcomes by treating each physical program as an executable hypothesis. In addition, we design a set of physics-control APIs to support more stable and complex motion behaviors. Extensive experiments demonstrate that PhysAgent produces more physically plausible videos, achieves better prompt alignment, and generalizes more effectively across diverse physical scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。