arXiv:2605.09269cs.CLcs.CV2026-05被引 1

让AI自己设计检查清单,精准判断图文回答差异。

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification

论文配图:DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
图 1 · 摘自论文原文
  • AI分两步走:先生成针对具体图像的验证清单,再逐项核对。
  • 在视觉奖励基准上,模型准确率提升22.6和18.8个百分点。
  • 适合需要高精度图文评估的研究者与开发者。

对齐多模态大语言模型需可靠奖励模型,但现有单步评估易因懒惰判断或语言先验而失准。虽然基于评分标准的评估可缓解文本任务中的偏差,但在多模态任务中受限于视觉推理复杂性。关键差异常依赖特定实例的视觉细节。鲁棒评估需动态生成能分离空间与事实差异的评分标准。为此,我们提出DeltaRubric,将多模态偏好评估重构为单一多模态大模型内的计划与执行过程。该方法分两步:首先作为‘分歧规划器’生成中立、实例相关的验证清单;随后转为‘清单验证器’,依据图像与问题执行自生成检查,得出最终有根据的判断。我们将DeltaRubric建模为多角色强化学习问题,联合优化规划与验证能力。在Qwen3-VL 4B和8B Instruct模型上验证,结果显著:在VL-RewardBench上,基线模型准确率分别提升+22.6(4B)和+18.8(8B),远超无评分标准的基线。结果表明,将评估分解为结构化可验证步骤,可实现更可靠、泛化更强的多模态奖励建模。

原文摘要 · Abstract (English)

Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exploiting language priors over fine-grained visual verification. While rubric-based evaluation mitigates these biases in text-only settings, extending it to multimodal tasks is bottlenecked by the complexity of visual reasoning. The critical differences between responses often depend on instance-specific visual details. Robust evaluation requires dynamically synthesizing rubrics that isolate spatial and factual discrepancies. To address this, we introduce $\textbf{DeltaRubric}$, an approach that reformulates multimodal preference evaluation as a plan-and-execute process within a single MLLM. DeltaRubric operates in two steps: acting first as a $\textit{Disagreement Planner}$, the model generates a neutral, instance-specific verification checklist. Transitioning into a $\textit{Checklist Verifier}$, it executes these self-generated checks against the image and question to produce the final grounded judgment. We formulate DeltaRubric as a multi-role reinforcement learning problem, jointly optimizing planning and verification capabilities. Validated on Qwen3-VL 4B and 8B Instruct models, DeltaRubric achieves solid empirical gains. For instance, On VL-RewardBench, it improves base model overall accuracy by $\textbf{+22.6}$ (4B) and $\textbf{+18.8}$ (8B) points, largely outperforming standard no-rubric baselines. The results demonstrate that decomposing evaluation into structured, verifiable steps leads to more reliable and generalizable multimodal reward modeling.

多模态评估奖励模型视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。