arXiv:2511.09067cs.CLcs.AI2025-11EMNLP被引 2

评测大模型在图文任务中的批评能力,发现其自我改进潜力与表现相关。

MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique

  • 构建多维度图文批评评测基准,覆盖8类任务500+具体任务
  • 用专家标注的参考答案指导GPT-4o评分,提升判断可靠性
  • 发现模型输出质量与批评能力正相关,不同维度批评难度不同

批评能力对模型自我优化和成为可靠AI助手至关重要。尽管语言模型的批评能力已广泛研究,但大型多模态模型(LMMs)在图像描述、视觉推理等任务中日益增强,其多模态批评能力却仍鲜有探索。本文提出MM-CRITIC,一个涵盖基础、修正与比较三维度的综合性评测基准,包含8个主要任务类型和超过500项具体任务,共收集4471个样本,覆盖多种不同规模的LMMs。为提升评估可靠性,我们引入专家制定的参考答案作为评分标准,指导GPT-4o进行响应标注与参考批评生成,作为可信判断的锚点。大量实验验证了MM-CRITIC的有效性,并对主流LMMs在多个维度下的批评能力进行了全面评估。进一步分析揭示关键洞见:响应质量与批评能力存在相关性,且不同评估维度的批评难度各异。代码已开源。

原文摘要 · Abstract (English)

The ability of critique is vital for models to self-improve and serve as reliable AI assistants. While extensively studied in language-only settings, multimodal critique of Large Multimodal Models (LMMs) remains underexplored despite their growing capabilities in tasks like captioning and visual reasoning. In this work, we introduce MM-CRITIC, a holistic benchmark for evaluating the critique ability of LMMs across multiple dimensions: basic, correction, and comparison. Covering 8 main task types and over 500 tasks, MM-CRITIC collects responses from various LMMs with different model sizes and is composed of 4471 samples. To enhance the evaluation reliability, we integrate expert-informed ground answers into scoring rubrics that guide GPT-4o in annotating responses and generating reference critiques, which serve as anchors for trustworthy judgments. Extensive experiments validate the effectiveness of MM-CRITIC and provide a comprehensive assessment of leading LMMs' critique capabilities under multiple dimensions. Further analysis reveals some key insights, including the correlation between response quality and critique, and varying critique difficulty across evaluation dimensions. Our code is available at https://github.com/MichealZeng0420/MM-Critic.

多模态模型评估批评能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。