让视觉语言模型在测试时像人一样思考偏好,实现可解释的图像排序。
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
- 用强化学习让VLM在测试时推理人类视觉偏好,基于软奖励策略。
- 在ImageReward和HPSv2数据集上分别达到64.9%和65.4%准确率。
- 无需大量标注,可解释性强,适合提升文生图模型的偏好优化。
视觉语言模型能否有效捕捉人类视觉偏好?本文通过在测试时引入基于DeepSeek R1和OpenAI O1启发的强化学习方法,训练VLM进行偏好推理。利用ImageReward和Human Preference Score v2(HPSv2)数据集,模型在ImageReward测试集(使用官方划分)上达到64.9%准确率,在HPSv2上达到65.4%(仅用约25%数据训练)。结果与传统编码器模型相当,同时具备透明推理和更强泛化能力。该方法不仅利用VLM的世界知识,还发挥其推理潜力,生成可解释的决策依据。实验表明,当前VLM能合理建模人类视觉偏好,提出高效软奖励策略用于图像排序,优于简单选择或打分方法。该能力使VLM可对任意尺寸、复杂度图像进行排序,有望显著提升视觉偏好优化效果。相比依赖大量标注的方案,本方法减少标注需求,提升奖励泛化性与可解释性,为文本到视觉模型的进一步优化提供重要里程碑。
原文摘要 · Abstract (English)
Can Visual Language Models (VLMs) effectively capture human visual preferences? This work addresses this question by training VLMs to think about preferences at test time, employing reinforcement learning methods inspired by DeepSeek R1 and OpenAI O1. Using datasets such as ImageReward and Human Preference Score v2 (HPSv2), our models achieve accuracies of 64.9% on the ImageReward test set (trained on ImageReward official split) and 65.4% on HPSv2 (trained on approximately 25% of its data). These results match traditional encoder-based models while providing transparent reasoning and enhanced generalization. This approach allows to use not only rich VLM world knowledge, but also its potential to think, yielding interpretable outcomes that help decision-making processes. By demonstrating that human visual preferences reasonable by current VLMs, we introduce efficient soft-reward strategies for image ranking, outperforming simplistic selection or scoring methods. This reasoning capability enables VLMs to rank arbitrary images-regardless of aspect ratio or complexity-thereby potentially amplifying the effectiveness of visual Preference Optimization. By reducing the need for extensive markup while improving reward generalization and explainability, our findings can be a strong mile-stone that will enhance text-to-vision models even further.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。