arXiv:2603.11665cs.CL2026-03ACL

用多任务强化学习让图文大模型更懂多场景评价

Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge

  • 通过多任务强化学习统一优化多个评估任务
  • 在跨任务一致性与人类偏好相关性上超越基线
  • 擅长处理分布外任务,适合复杂评测场景

图文大模型(MLLM)因其在多种视觉任务中与人类判断高度对齐,被广泛用作MLLM-as-a-Judge。然而,现有评判模型多针对单任务优化,在多样化场景下泛化能力不足,这限制了其可靠评估性能。为此,我们提出多任务强化学习框架MT-RL-Judge,通过联合优化多个任务,利用强化学习的泛化能力提升评判模型表现。实验表明,该方法在判断一致性与人类偏好相关性上优于多个强基线模型;同时在分布外任务上仍保持稳健表现,验证了其有效性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have been widely adopted as MLLM-as-a-Judges due to their strong alignment with human judgment across various visual tasks. However, most existing judge models are optimized for single-task scenarios and struggle to generalize to diverse contexts, which is a critical requirement for reliable evaluation. To address this limitation, we propose Multi-Task Reinforcement Learning for MLLM-as-a-Judge (MT-RL-Judge), a framework that jointly optimizes the judge model across multiple tasks, leveraging the generalization capabilities of RL. Experimental results against several strong baselines demonstrate that MT-RL-Judge outperforms strong baselines in both judgment consistency and correlation with human preferences. Furthermore, our approach exhibits robust generalization on out-of-distribution tasks, further validating its effectiveness.

多任务学习强化学习图文模型自动评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。