让视频生成遵循物理规律,用代码控制运动轨迹。
MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
- 用多智能体大模型+物理引擎生成可复现的物理准确视频
- 在1275个标注视频上验证,现有模型物理准确性普遍不足
- 适合关注物理一致性与可控生成的研究者和开发者
尽管文本到视频(T2V)生成在逼真度上取得显著进展,但生成符合意图且严格遵守物理规律的视频仍是核心挑战。本文系统研究牛顿力学控制下的文本到视频生成与评估,强调物理精确性与运动连贯性。提出MoReGen,一种基于多智能体大模型、物理模拟器与渲染器的运动感知、物理驱动的T2V框架,从文本提示生成可复现、物理准确的代码域视频。为定量评估物理有效性,提出物体轨迹对应性作为直接评价指标,并构建了包含1275个由人类标注的视频的MoReSet基准数据集,涵盖九类牛顿现象,包含场景描述、时空关系与真实轨迹。基于此,对现有T2V模型进行实验,通过我们的MoRe指标与现有物理评估方法对比,结果表明当前顶尖模型难以保持物理有效性,而MoReGen为实现物理一致的视频合成提供了原则性方向。
原文摘要 · Abstract (English)
While text-to-video (T2V) generation has achieved remarkable progress in photorealism, generating intent-aligned videos that faithfully obey physics principles remains a core challenge. In this work, we systematically study Newtonian motion-controlled text-to-video generation and evaluation, emphasizing physical precision and motion coherence. We introduce MoReGen, a motion-aware, physics-grounded T2V framework that integrates multi-agent LLMs, physics simulators, and renderers to generate reproducible, physically accurate videos from text prompts in the code domain. To quantitatively assess physical validity, we propose object-trajectory correspondence as a direct evaluation metric and present MoReSet, a benchmark of 1,275 human-annotated videos spanning nine classes of Newtonian phenomena with scene descriptions, spatiotemporal relations, and ground-truth trajectories. Using MoReSet, we conduct experiments on existing T2V models, evaluating their physical validity through both our MoRe metrics and existing physics-based evaluators. Our results reveal that state-of-the-art models struggle to maintain physical validity, while MoReGen establishes a principled direction toward physically coherent video synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。