arXiv:2604.13029cs.CVcs.AI2026-04被引 5

用细粒度评分标准提升视觉任务的偏好优化效果

Visual Preference Optimization with Rubric Rewards

论文配图:Visual Preference Optimization with Rubric Rewards
图 1 · 摘自论文原文
  • 为每张图-指令对设计专属评分清单,实现精准反馈
  • 在下游任务中将平均分提升至82.69,显著优于传统方法
  • 适合需要高精度视觉评估的生成模型训练场景

直接偏好优化(DPO)的效果依赖于能反映多模态任务质量差异的偏好数据。现有方法常依赖离策略扰动或粗粒度结果信号,难以支持细粒度视觉推理。本文提出rDPO,基于实例特定的评分标准进行偏好优化。针对每个图像-指令对,构建包含核心与附加标准的检查表式评分清单,用于评估任意策略的响应。评分标准池在离线阶段建立,供在线策略数据构建时重复使用。在公开的奖励建模基准上,基于评分标准的提示使30B-A3B判别器表现接近GPT-5.4;在公开下游基准上,基于评分标准的过滤使宏平均分升至82.69,而基于结果的过滤则从81.14降至75.82。在综合性基准上的可扩展性评估中,rDPO达到61.01,明显优于风格约束基线(52.36),并超越基础模型的59.48。结果表明,结合在线策略数据构建与实例级标准反馈,可显著提升视觉偏好优化效果。

原文摘要 · Abstract (English)

The effectiveness of Direct Preference Optimization (DPO) depends on preference data that reflect the quality differences that matter in multimodal tasks. Existing pipelines often rely on off-policy perturbations or coarse outcome-based signals, which are not well suited to fine-grained visual reasoning. We propose rDPO, a preference optimization framework based on instance-specific rubrics. For each image-instruction pair, we create a checklist-style rubric of essential and additional criteria to score responses from any possible policies. The instruction-rubric pool is built offline and reused during the construction of on-policy data. On public reward modeling benchmarks, rubric-based prompting massively improves a 30B-A3B judge and brings it close to GPT-5.4. On public downstream benchmarks, rubric-based filtering raises the macro average to 82.69, whereas outcome-based filtering drops it to 75.82 from 81.14. When evaluating scalability on a comprehensive benchmark, rDPO achieves 61.01, markedly outperforming the style-constrained baseline (52.36) and surpassing the 59.48 base model. Together, these results show that visual preference optimization benefits from combining on-policy data construction with instance-specific criterion-level feedback.

偏好优化视觉推理评分标准生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。