arXiv:2605.22818cs.CV2026-05被引 1

让视频生成懂因果,运动控制更自然

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

论文配图:MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
图 1 · 摘自论文原文
  • 先推理后生成,用视觉语言模型补全运动逻辑
  • 在新基准上显著提升动作合理性,人类评测更偏好
  • 适合需要真实交互的视频生成场景

当前运动控制的图像到视频生成模型依赖稀疏、不准确且因果不完整的轨迹,常导致不自然或不合理的结果,尤其忽略次要因果效应。为此,我们提出MotiMotion,将运动控制重构为‘推理-生成’问题。通过无需训练的视觉语言模型(VLM),对主轨迹的图像空间坐标进行修正,并推断合理的次级运动。进一步提出置信度感知的控制机制,根据输入可靠性调节引导强度:高置信时精准跟随,低置信时利用模型内在生成先验修正错误。为系统评估,我们构建了新基准MotiBench,包含由运动触发新事件的交互中心场景。基于VLM和真人评测均表明,MotiMotion生成的视频在物体行为与交互合理性上显著优于现有方法。

原文摘要 · Abstract (English)

Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often yields unnatural or implausible outcomes, especially by missing secondary causal consequences. To address this, we introduce MotiMotion, a novel framework that reformulates motion control as a reasoning-then-generation problem. To encourage causally grounded and commonsense-consistent interactions, we leverage a training-free vision-language reasoner to refine image-space coordinates of primary trajectories and to hallucinate plausible secondary motions. To further improve motion naturalness, we propose a confidence-aware control scheme that modulates guidance strength, enabling the model to closely follow high-confidence plans while correcting artifacts under low-confidence inputs with its internal generative priors. To support systematic evaluation, we curate a new image-to-video benchmark, MotiBench, consisting of interaction-centric scenes where new events are triggered by motion. Both VLM-based evaluation and a human study on MotiBench demonstrate that MotiMotion produces videos with more plausible object behaviors and interaction, and is preferred over existing approaches.

视频生成运动控制因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。