arXiv:2412.16153cs.CVcs.AI2024-12CVPR被引 10

用运动聚焦损失提升文本引导图像动画的对齐效果

MotiF: Making Text Count in Image Animation with Motion Focal Loss

  • 根据运动热图加权损失,让模型更关注动态区域
  • 在新基准上优于9个开源模型,平均偏好率达72%
  • 适合需要精准文本-动作对齐的视频生成研究者

文本-图像到视频(TI2V)生成旨在根据图像和文本描述生成视频,也称为文本引导图像动画。现有方法在生成与文本提示一致的视频时表现不佳,尤其在涉及运动描述时。为此,我们提出MotiF,一种简单而有效的方法,通过将模型学习聚焦于运动更显著的区域,提升文本对齐与运动生成质量。利用光流生成运动热图,并按运动强度加权损失。该改进目标显著提升性能,且可与使用运动先验作为输入的方法互补。此外,由于缺乏多样化的TI2V评估基准,我们构建了TI2V Bench,包含320对图像-文本数据用于稳健评估。我们提出人类评价协议,要求标注者在两段视频中选择整体偏好并给出理由。在TI2V Bench上的全面评估显示,MotiF超越九个开源模型,平均偏好率达到72%。TI2V Bench及额外结果已公开于https://wang-sj16.github.io/motif/。

原文摘要 · Abstract (English)

Text-Image-to-Video (TI2V) generation aims to generate a video from an image following a text description, which is also referred to as text-guided image animation. Most existing methods struggle to generate videos that align well with the text prompts, particularly when motion is specified. To overcome this limitation, we introduce MotiF, a simple yet effective approach that directs the model's learning to the regions with more motion, thereby improving the text alignment and motion generation. We use optical flow to generate a motion heatmap and weight the loss according to the intensity of the motion. This modified objective leads to noticeable improvements and complements existing methods that utilize motion priors as model inputs. Additionally, due to the lack of a diverse benchmark for evaluating TI2V generation, we propose TI2V Bench, a dataset consists of 320 image-text pairs for robust evaluation. We present a human evaluation protocol that asks the annotators to select an overall preference between two videos followed by their justifications. Through a comprehensive evaluation on TI2V Bench, MotiF outperforms nine open-sourced models, achieving an average preference of 72%. The TI2V Bench and additional results are released in https://wang-sj16.github.io/motif/.

视频生成运动对齐文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。