提出首个视频生成中运动归因框架,可精准识别影响动作的关键数据。
Motion Attribution for Video Generation
- 基于梯度的运动归因方法,分离动态与静态特征,高效计算数据影响。
- 使用精选高影响力数据微调后,动作流畅度与动态性提升,人工偏好胜率达74.1%。
- 适用于文本到视频模型的数据筛选,提升生成结果的时间一致性与物理合理性。
尽管视频生成模型进展迅速,但数据对运动的影响机制仍不清晰。我们提出Motive(MOTIon attribution for Video gEneration),一个面向运动、基于梯度的数据归因框架,可扩展至现代大规模高质量视频数据集与模型。通过运动加权损失掩码,Motive将时间动态与静态外观分离,实现高效的运动特异性影响计算。在文本到视频模型上,Motive识别出显著影响运动的微调片段,并指导数据筛选,提升了时间一致性和物理合理性。采用Motive筛选的高影响力数据后,本方法在VBench上使动作平滑度和动态程度均获提升,人工偏好胜率达到74.1%,优于预训练基线模型。据我们所知,这是首个针对视频生成模型中运动而非视觉外观进行归因,并用于数据筛选的框架。
原文摘要 · Abstract (English)
Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework that scales to modern, large, high-quality video datasets and models. We use this to study which fine-tuning clips improve or degrade temporal dynamics. Motive isolates temporal dynamics from static appearance via motion-weighted loss masks, yielding efficient and scalable motion-specific influence computation. On text-to-video models, Motive identifies clips that strongly affect motion and guides data curation that improves temporal consistency and physical plausibility. With Motive-selected high-influence data, our method improves both motion smoothness and dynamic degree on VBench, achieving a 74.1% human preference win rate compared with the pretrained base model. To our knowledge, this is the first framework to attribute motion rather than visual appearance in video generative models and to use it to curate fine-tuning data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。