用交互图控视频生成,让多对象运动更灵活精准。
GraphVid: Interactive Graph-Controllable Video Generation

- 以结构化交互图作为控制输入,替代繁琐轨迹绘制。
- 在少量数据和参数下,视频质量显著优于现有方法。
- 适合需要精细多主体控制的视频创作与研究者使用。
可控视频生成因难以通过文本提示或运动控制精确描述多对象交互而面临挑战。传统基于轨迹的控制需用户绘制多个物体的精确路径,复杂场景下效率低且易受遮挡影响。为此,我们提出GraphVid,一种基于图条件的图像到视频生成模型,通过结构化交互图实现交互式控制。同时构建了大规模以互动为核心的GraphVid-Bench数据集,包含结构化关系标注,用于训练具备交互感知能力的视频生成模型。尽管训练数据和可训练参数远少于先前运动控制方法,GraphVid在可控性和视频质量上均表现优异:相比Motion-I2V,FID降低39.9%,FVD降低37.6%,PSNR从9.87提升至15.98,SSIM从0.38提升至0.61。结果表明,结构化语义接口是可控视频生成的强大范式。
原文摘要 · Abstract (English)
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce $\textbf{GraphVid}$, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate $\textbf{GraphVid-Bench}$, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。