通过运动场代理实现文本引导视频生成的精细动作控制
MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent
- 将文本中的物体与镜头运动转为轨迹和相机参数
- 在VBench上相机运动控制指标显著提升
- 适合需要精准镜头与物体运动控制的场景
我们提出MotionAgent,实现文本引导图像到视频生成中的细粒度运动控制。核心是运动场代理,将文本提示中的运动信息转化为显式的运动场,提供灵活精确的运动引导。具体而言,该代理提取文本描述的物体运动和相机运动,并分别转换为物体轨迹和相机外参。一个解析光流组合模块在三维空间中融合这些运动表征,并投影为统一光流。光流适配器利用该光流控制基础图像到视频扩散模型,生成精细可控视频。VBench上的视频-文本相机运动指标显著提升,表明方法对相机运动具有精确控制能力。我们构建了VBench的一个子集,用于评估文本与生成视频间运动信息的一致性,在运动生成准确性上优于其他先进模型。
原文摘要 · Abstract (English)
We propose MotionAgent, enabling fine-grained motion control for text-guided image-to-video generation. The key technique is the motion field agent that converts motion information in text prompts into explicit motion fields, providing flexible and precise motion guidance. Specifically, the agent extracts the object movement and camera motion described in the text and converts them into object trajectories and camera extrinsics, respectively. An analytical optical flow composition module integrates these motion representations in 3D space and projects them into a unified optical flow. An optical flow adapter takes the flow to control the base image-to-video diffusion model for generating fine-grained controlled videos. The significant improvement in the Video-Text Camera Motion metrics on VBench indicates that our method achieves precise control over camera motion. We construct a subset of VBench to evaluate the alignment of motion information in the text and the generated video, outperforming other advanced models on motion generation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。