MagicMotion实现从密集到稀疏轨迹的可控视频生成,支持多物体精准运动控制。
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
- 通过掩码、边界框和稀疏框三种轨迹条件实现灵活控制
- 在多物体场景下保持运动一致性与高质量视觉效果
- 配套数据集与评测基准,适合研究可控视频生成的学者
近期视频生成技术在视觉质量与时间连贯性上取得显著进步。在此基础上,轨迹可控视频生成通过显式定义空间路径实现物体运动精确控制。然而,现有方法在复杂运动和多物体控制中表现不佳,存在轨迹偏离、对象一致性差及画质下降等问题。此外,这些方法仅支持单一轨迹格式,适用性受限,且缺乏专门针对轨迹可控视频生成的公开数据集与评估基准。为此,我们提出MagicMotion,一种新型图像到视频生成框架,支持从密集到稀疏的三类轨迹条件:掩码、边界框与稀疏框。给定输入图像与轨迹,MagicMotion可无缝沿指定路径动画化物体,同时保持对象一致性和视觉质量。我们还构建了大规模轨迹可控视频数据集MagicData,以及自动化标注与过滤流程。此外,提出MagicBench评测基准,系统评估不同物体数量下的视频质量与轨迹控制精度。大量实验表明,MagicMotion在各项指标上均优于先前方法。项目页面已公开:https://quanhaol.github.io/magicmotion-site。
原文摘要 · Abstract (English)
Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly defined spatial paths. However, existing methods struggle with complex object movements and multi-object motion control, resulting in imprecise trajectory adherence, poor object consistency, and compromised visual quality. Furthermore, these methods only support trajectory control in a single format, limiting their applicability in diverse scenarios. Additionally, there is no publicly available dataset or benchmark specifically tailored for trajectory-controllable video generation, hindering robust training and systematic evaluation. To address these challenges, we introduce MagicMotion, a novel image-to-video generation framework that enables trajectory control through three levels of conditions from dense to sparse: masks, bounding boxes, and sparse boxes. Given an input image and trajectories, MagicMotion seamlessly animates objects along defined trajectories while maintaining object consistency and visual quality. Furthermore, we present MagicData, a large-scale trajectory-controlled video dataset, along with an automated pipeline for annotation and filtering. We also introduce MagicBench, a comprehensive benchmark that assesses both video quality and trajectory control accuracy across different numbers of objects. Extensive experiments demonstrate that MagicMotion outperforms previous methods across various metrics. Our project page are publicly available at https://quanhaol.github.io/magicmotion-site.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。