提出可分离物体与相机运动强度的新估计算法,提升图像转视频质量。
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
- 通过对比学习自动估算视频中物体与相机的独立运动强度
- 在大规模真实视频上实现稳定运动估计,支持高质量图像转视频生成
- 适用于数据处理与视频生成训练,可作为通用增强模块
图像转视频(I2V)生成通常以静态图像为条件,近期通过引入运动强度作为额外控制信号得到增强。这类运动感知模型能生成多样化的运动模式,但缺乏可靠的运动估计算法来训练大规模真实视频数据集。传统指标如SSIM或光流难以泛化至任意视频,而人工标注抽象运动强度也极困难。此外,运动强度需同时反映局部物体运动与全局相机运动,此前尚未被研究。本文提出一种新型运动估计算法,可测量视频中物体与相机的解耦运动强度。利用随机配对视频的对比学习,区分运动强度更高的视频,该范式易于标注且便于扩展,实现稳定的运动估计性能。基于此,我们构建了新I2V模型MotionStone。实验表明,所提运动估计算法稳定,MotionStone在I2V生成任务上达到当前最优表现。该解耦运动估计算法可作为通用插件,用于数据处理与视频生成训练。
原文摘要 · Abstract (English)
The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such models on large-scale video set in the wild. Traditional metrics, e.g., SSIM or optical flow, are hard to generalize to arbitrary videos, while, it is very tough for human annotators to label the abstract motion intensity neither. Furthermore, the motion intensity shall reveal both local object motion and global camera movement, which has not been studied before. This paper addresses the challenge with a new motion estimator, capable of measuring the decoupled motion intensities of objects and cameras in video. We leverage the contrastive learning on randomly paired videos and distinguish the video with greater motion intensity. Such a paradigm is friendly for annotation and easy to scale up to achieve stable performance on motion estimation. We then present a new I2V model, named MotionStone, developed with the decoupled motion estimator. Experimental results demonstrate the stability of the proposed motion estimator and the state-of-the-art performance of MotionStone on I2V generation. These advantages warrant the decoupled motion estimator to serve as a general plug-in enhancer for both data processing and video generation training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。