arXiv:2603.07990cs.LG2026-03被引 1

用视觉证据链提升多模态判断准确率,小模型超越大模型。

MJ1: Multimodal Judgment via Grounded Verification

  • 构建视觉验证链,强制模型基于图像证据做决策。
  • 训练后在MMRB2上达77.0%准确率,优于千亿级模型Gemini-3-Pro。
  • 无需额外训练即可提升基线模型性能,适合多模态评估场景。

多模态判别器常无法基于视觉证据做出判断。本文提出MJ1,一种通过强化学习训练的多模态判别器,采用结构化视觉验证链(观察→陈述→验证→评估→打分)并引入反事实一致性奖励以惩罚立场偏差。即使未经训练,该机制也能使基线模型在MMRB2数据集上图像编辑任务提升+3.8分,多模态推理任务提升+1.7分。训练后,仅30亿活跃参数的MJ1在MMRB2上达到77.0%准确率,超越数量级更大的Gemini-3-Pro模型。结果表明,基于视觉接地与一致性训练的方法显著提升多模态判断能力,且无需增加模型规模。

原文摘要 · Abstract (English)

Multimodal judges struggle to ground decisions in visual evidence. We present MJ1, a multimodal judge trained with reinforcement learning that enforces visual grounding through a structured grounded verification chain (observations $\rightarrow$ claims $\rightarrow$ verification $\rightarrow$ evaluation $\rightarrow$ scoring) and a counterfactual consistency reward that penalizes position bias. Even without training, our mechanism improves base-model accuracy on MMRB2 by +3.8 points on Image Editing and +1.7 on Multimodal Reasoning. After training, MJ1, with only 3B active parameters, achieves 77.0% accuracy on MMRB2 and surpasses orders-of-magnitude larger models like Gemini-3-Pro. These results show that grounded verification and consistency-based training substantially improve multimodal judgment without increasing model scale.

多模态判断视觉验证小模型大能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。