arXiv:2412.04814cs.CV2024-12被引 61

用人类反馈提升文本生成视频的准确性和偏好对齐

LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

  • 构建包含评分与理由的1万条人类标注数据集
  • 训练可捕捉人类偏好的奖励模型,提升视频对齐度
  • 在CogVideoX上微调后超越5B大模型性能

文本到视频生成模型虽有显著进展,但在准确反映文本描述和符合人类偏好方面仍不足。现有方法依赖人工评分的视频质量评估模型,却忽略评价背后的理由,难以捕捉细微偏好。本文提出首个利用人类反馈对齐文本生成视频模型的方法LiFT。首先构建包含约10,000条标注的LiFT-HRA数据集,每条含评分与理由;基于此训练奖励模型LiFT-Critic,学习人类判断的奖励函数;最后通过最大化奖励加权似然对模型进行对齐优化。以CogVideoX-2B为例,微调后模型在全部16项指标上优于原始的CogVideoX-5B,验证了人类反馈在提升生成质量与对齐度方面的潜力。

原文摘要 · Abstract (English)

Recent advances in text-to-video (T2V) generative models have shown impressive capabilities. However, these models are still inadequate in aligning synthesized videos with human preferences (e.g., accurately reflecting text descriptions), which is particularly difficult to address, as human preferences are subjective and challenging to formalize as objective functions. Existing studies train video quality assessment models that rely on human-annotated ratings for video evaluation but overlook the reasoning behind evaluations, limiting their ability to capture nuanced human criteria. Moreover, aligning T2V model using video-based human feedback remains unexplored. Therefore, this paper proposes LiFT, the first method designed to leverage human feedback for T2V model alignment. Specifically, we first construct a Human Rating Annotation dataset, LiFT-HRA, consisting of approximately 10k human annotations, each including a score and its corresponding rationale. Based on this, we train a reward model LiFT-Critic to learn reward function effectively, which serves as a proxy for human judgment, measuring the alignment between given videos and human expectations. Lastly, we leverage the learned reward function to align the T2V model by maximizing the reward-weighted likelihood. As a case study, we apply our pipeline to CogVideoX-2B, showing that the fine-tuned model outperforms the CogVideoX-5B across all 16 metrics, highlighting the potential of human feedback in improving the alignment and quality of synthesized videos.

文本生成视频人类反馈奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。