首个评估多模态模型在多元标准下判断能力的基准测试
Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- 构建包含多标准人类标注的挑战性数据集,评估模型对复杂评价标准的遵循能力
- 发现商用模型在开放生成任务中仍难保持多元标准一致性,开源模型更弱
- 提出新指标衡量标准切换灵活性与偏好冲突识别,推动可调控的AI评估发展
大型多模态模型(LMMs)因其强指令跟随能力和与人类偏好的一致性,正被广泛用作多模态评估系统的裁判。然而,其在遵循多样、细粒度评价标准方面的能力仍不明确。本文构建了Multi-Crit——一个用于评估多模态裁判在遵循多元标准及生成可靠标准级判断方面的基准。该基准涵盖开放生成与可验证推理任务,通过严谨的数据筛选流程收集具有挑战性的响应对,并附带多标准人类标注。此外,引入三种新指标,系统评估多元遵循性、标准切换灵活性及准则级偏好冲突识别能力。对25个LMM的综合分析表明:1)商用模型在开放生成任务中仍难以维持多元标准的一致性;2)开源模型在灵活遵循多样化标准方面显著落后;3)以整体判断信号进行批评微调虽增强视觉定位能力,但无法泛化至多元准则级判断。进一步分析揭示了推理微调、测试时扩展及开源与商用模型间边界一致性等问题,为构建可靠、可调控的多模态评估体系奠定基础。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) are increasingly adopted as judges in multimodal evaluation systems due to their strong instruction following and consistency with human preferences. However, their ability to follow diverse, fine-grained evaluation criteria remains underexplored. We develop Multi-Crit, a benchmark for evaluating multimodal judges on their capacity to follow pluralistic criteria and produce reliable criterion-level judgments. Covering both open-ended generation and verifiable reasoning tasks, Multi-Crit is built through a rigorous data curation pipeline that gathers challenging response pairs with multi-criterion human annotations. It further introduces three novel metrics for systematically assessing pluralistic adherence, criterion-switching flexibility, and the ability to recognize criterion-level preference conflicts. Comprehensive analysis of 25 LMMs reveals that 1) proprietary models still struggle to maintain consistent adherence to pluralistic criteria--especially in open-ended evaluation; 2) open-source models lag further behind in flexibly following diverse criteria; and 3) critic fine-tuning with holistic judgment signals enhances visual grounding but fails to generalize to pluralistic criterion-level judgment. Additional analyses on reasoning fine-tuning, test-time scaling, and boundary consistency between open-source and proprietary models further probe the limits of current multimodal judges. As a pioneering study, Multi-Crit lays the foundation for building reliable and steerable multimodal AI evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。