arXiv:2506.22832cs.CVcs.AI2025-06被引 2

让模型推理时考虑独立判断者意见,提升视觉偏好对齐效果。

Listener-Rewarded Thinking in VLMs for Image Preferences

  • 引入独立监听模型校准推理过程,生成更可信的奖励信号。
  • 在ImageReward上达67.4%准确率,分布外性能提升最多6%。
  • 适合需要高可信推理的生成模型对齐任务,如图像/视频生成。

训练鲁棒且泛化的视觉偏好奖励模型对对齐文本到图像和文本到视频生成模型与人类意图至关重要。然而,现有奖励模型泛化能力差,监督微调易导致记忆,需复杂标注流程。尽管强化学习(特别是分组相对策略优化,GRPO)提升了泛化性,我们发现关键缺陷:当模型推理路径与独立冻结的视觉语言模型(“监听者”)评估结果矛盾时,推理准确率显著下降。为此,我们提出监听者增强型GRPO框架。监听者重新评估推理链,提供密集、校准的置信度评分,塑造强化学习奖励信号。这促使推理模型不仅回答正确,且生成能说服独立模型的解释。该监听者驱动的奖励机制在ImageReward基准上取得67.4%最佳准确率,大幅改善大规模人类偏好数据集(120万投票)上的分布外性能(最高+6%),并减少推理矛盾。结果表明,监听者奖励为对齐视觉语言模型与微妙人类偏好提供了可扩展、数据高效路径。推理模型将公开发布于:https://huggingface.co/alexgambashidze/qwen2.5vl_image_preference_reasoner。

原文摘要 · Abstract (English)

Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However, current reward models often fail to generalize, and supervised fine-tuning leads to memorization, demanding complex annotation pipelines. While reinforcement learning (RL), specifically Group Relative Policy Optimization (GRPO), improves generalization, we uncover a key failure mode: a significant drop in reasoning accuracy occurs when a model's reasoning trace contradicts that of an independent, frozen vision-language model ("listener") evaluating the same output. To address this, we introduce a listener-augmented GRPO framework. Here, the listener re-evaluates the reasoner's chain-of-thought to provide a dense, calibrated confidence score, shaping the RL reward signal. This encourages the reasoner not only to answer correctly, but to produce explanations that are persuasive to an independent model. Our listener-shaped reward scheme achieves best accuracy on the ImageReward benchmark (67.4%), significantly improves out-of-distribution (OOD) performance on a large-scale human preference dataset (1.2M votes, up to +6% over naive reasoner), and reduces reasoning contradictions compared to strong GRPO and SFT baselines. These results demonstrate that listener-based rewards provide a scalable, data-efficient path to aligning vision-language models with nuanced human preferences. We will release our reasoning model here: https://huggingface.co/alexgambashidze/qwen2.5vl_image_preference_reasoner.

视觉偏好强化学习推理对齐奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。