让视频生成学会物理常识,用智能规划解决生成结果违背物理规律的问题
NEWTON: Agentic Planning for Physically Grounded Video Generation

- 构建代理式规划系统,通过关键帧生成、科学计算和提示优化构建物理条件
- 在VideoPhy-2上将联合准确率提升至37.4%,显著优于现有模型
- 无需修改生成器,适合需要高物理真实性的视频生成任务
视频生成模型虽视觉效果出色,但普遍存在违反物理常识的问题——在VideoPhy-2数据集上,最佳模型的联合准确率仅达32.6%。我们发现根本瓶颈在于文本提示对物理世界的损失性压缩:关键动力学参数被遗漏,模型规模再大也无法弥补缺失信息。基于此诊断,提出物理条件必须满足充分性、动态性与可验证性三个特性,并指出现有方法均不满足全部三项。为此提出NEWTON框架,将视频生成从输出环节降级为智能体工具箱中的一个动作:由学习到的规划器协调物理感知工具(关键帧生成、科学计算、提示优化)构建丰富条件,验证器闭环反馈实现迭代重规划。规划器是唯一可训练组件,通过流式在线策略优化(Flow-GRPO)在多轮交互中持续改进。在VideoPhy-2上,NEWTON使LTX-Video的联合准确率从21.4%提升至29.7%,Veo-3.1从30.7%提升至37.4%,且未修改任何生成模型。
原文摘要 · Abstract (English)
Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves only 32.6% joint accuracy. We identify a specification bottleneck: text prompts are lossy compression of the physical world, omitting the parameters that fully determine dynamics, and no amount of model scaling can recover what was never specified. From this diagnosis we derive three properties that physics conditioning must satisfy -- sufficiency, dynamism, and verifiability -- and show that no existing approach satisfies all three. We present NEWTON, in which video generation is demoted from the system output to one action inside an agent's toolbox: a learned planner orchestrates physics-aware tools (keyframe generation, scientific computation, prompt refinement) to construct rich conditioning, and a verifier closes the loop for iterative re-planning. The planner is the sole trainable component, optimized on-policy via Flow-GRPO inside the live multi-turn loop. On VideoPhy-2, NEWTON improves joint accuracy from 21.4% to 29.7% on LTX-Video and from 30.7% to 37.4% on Veo-3.1, without modifying either generator. Our project page: https://Newton026.github.io/newton
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。