arXiv:2412.02617cs.LGcs.AI2024-12被引 47

用AI反馈提升文本生成视频中的物体动态真实感

Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

  • 用视觉语言模型提供针对物体动态的感知反馈
  • 在多物体交互和物体下落场景中显著提升真实感
  • 适合关注视频生成物理合理性的研究者

大型文本到视频模型在众多下游应用中潜力巨大,但难以准确刻画动态物体交互,常导致运动不真实、违背现实物理规律。受大语言模型启发,我们探索利用外部反馈来对齐生成结果与预期目标。核心问题是:何种反馈搭配何种自改进算法,最能克服运动错位与交互失真?我们指出,文本到视频模型的离线强化学习微调算法可统一为一个概率目标,理论上无算法主导方法,关键在于奖励信号与数据特性。人类反馈虽精准但难扩展,而视觉语言模型可如人般感知视频场景。因此,我们提出使用视觉语言模型提供专用于物体动态的感知反馈。实验表明,相比主流视频质量指标,采用二值化AI反馈能显著提升交互场景的视频质量,经AI、人工及质量度量评估均验证有效。尤其在多物体复杂交互与真实下落模拟中收益显著。

原文摘要 · Abstract (English)

Large text-to-video models hold immense potential for a wide range of downstream applications. However, they struggle to accurately depict dynamic object interactions, often resulting in unrealistic movements and frequent violations of real-world physics. One solution inspired by large language models is to align generated outputs with desired outcomes using external feedback. In this work, we investigate the use of feedback to enhance the quality of object dynamics in text-to-video models. We aim to answer a critical question: what types of feedback, paired with which specific self-improvement algorithms, can most effectively overcome movement misalignment and realistic object interactions? We first point out that offline RL-finetuning algorithms for text-to-video models can be equivalent as derived from a unified probabilistic objective. This perspective highlights that there is no algorithmically dominant method in principle; rather, we should care about the property of reward and data. While human feedback is less scalable, vision-language models could notice the video scenes as humans do. We then propose leveraging vision-language models to provide perceptual feedback specifically tailored to object dynamics in videos. Compared to popular video quality metrics measuring alignment or dynamics, the experiments demonstrate that our approach with binary AI feedback drives the most significant improvements in the quality of interaction scenes in video, as confirmed by AI, human, and quality metric evaluations. Notably, we observe substantial gains when using signals from vision language models, particularly in scenarios involving complex interactions between multiple objects and realistic depictions of objects falling.

文本生成视频物体交互AI反馈视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。