arXiv:2412.21059cs.CV2024-12AAAI被引 156

让生成图像视频更符合人类偏好,还能解释为什么好。

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

论文配图:VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
图 1 · 摘自论文原文
  • 分层评估+线性加权,可解释地学习人类视觉偏好
  • 在图片视频生成中,预测准确率比现有模型高17.2%
  • 适合需要高质量、可解释生成结果的研究者

视觉生成模型在合成逼真图像和视频方面取得了显著进展,但使其输出与人类在关键维度上的偏好对齐仍是持续挑战。尽管基于人类反馈的强化学习在偏好对齐方面具有潜力,但现有的视觉生成奖励模型存在评分黑箱化、缺乏可解释性,可能导致意外偏差。我们提出 VisionReward,一个通用框架,用于学习图像和视频生成中的人类视觉偏好。具体而言,我们采用分层视觉评估框架捕捉细粒度人类偏好,并利用线性加权实现可解释的偏好学习。此外,我们在使用 VisionReward 作为视觉生成偏好优化中的奖励模型时,提出了多维一致性策略。实验表明,VisionReward 在机器指标和人工评估上均显著优于现有图像和视频奖励模型。值得注意的是,VisionReward 在偏好预测准确率上比 VideoScore 提高 17.2%,使用 VisionReward 的文本到视频模型相比使用 VideoScore 的模型,在配对胜率上高出 31.6%。所有代码和数据集已开源至 https://github.com/THUDM/VisionReward。

原文摘要 · Abstract (English)

Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement learning from human feedback offers promise for preference alignment, existing reward models for visual generation face limitations, including black-box scoring without interpretability and potentially resultant unexpected biases. We present VisionReward, a general framework for learning human visual preferences in both image and video generation. Specifically, we employ a hierarchical visual assessment framework to capture fine-grained human preferences, and leverages linear weighting to enable interpretable preference learning. Furthermore, we propose a multi-dimensional consistent strategy when using VisionReward as a reward model during preference optimization for visual generation. Experiments show that VisionReward can significantly outperform existing image and video reward models on both machine metrics and human evaluation. Notably, VisionReward surpasses VideoScore by 17.2% in preference prediction accuracy, and text-to-video models with VisionReward achieve a 31.6% higher pairwise win rate compared to the same models using VideoScore. All code and datasets are provided at https://github.com/THUDM/VisionReward.

视觉生成偏好学习可解释性视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。