评测大模型是否真懂图像,而非靠套路答题。
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
- 构建5500条弱监督数据+900条人工标注的多模态推理题
- 15个模型在答案准确、推理逻辑、视觉对齐三方面被严格评估
- 新增评分指标,发现模型常错但不自知,适合研究视觉理解的学者
大视觉语言模型在视觉问答和多模态推理上表现优异,但其是否真正进行基于图像的合理推理仍存疑。本文提出MagiC,一个全面的基准测试,用于评估模型的具身多模态认知能力,不仅关注答案准确性,还考察逐步推理质量及其与视觉证据的一致性。该基准包含约5,500条由强模型输出生成的弱监督问答样本,以及900条人工精标样本,含答案、推理链和边界框标注。我们评估了15个参数量从7B到70B的视觉语言模型,涵盖四个维度:最终答案正确性、推理有效性、视觉对齐度及自我修正能力。此外,诊断设置用于测试模型在对抗性视觉干扰下的鲁棒性,并评估其内省纠错能力。引入新指标如MagiScore和StepSense,通过全面分析揭示当前方法在具身视觉推理中的关键局限与改进空间。
原文摘要 · Abstract (English)
Recent advances in large vision-language models have led to impressive performance in visual question answering and multimodal reasoning. However, it remains unclear whether these models genuinely perform grounded visual reasoning or rely on superficial patterns and dataset biases. In this work, we introduce MagiC, a comprehensive benchmark designed to evaluate grounded multimodal cognition, assessing not only answer accuracy but also the quality of step-by-step reasoning and its alignment with relevant visual evidence. Our benchmark includes approximately 5,500 weakly supervised QA examples generated from strong model outputs and 900 human-curated examples with fine-grained annotations, including answers, rationales, and bounding box groundings. We evaluate 15 vision-language models ranging from 7B to 70B parameters across four dimensions: final answer correctness, reasoning validity, grounding fidelity, and self-correction ability. MagiC further includes diagnostic settings to probe model robustness under adversarial visual cues and assess their capacity for introspective error correction. We introduce new metrics such as MagiScore and StepSense, and provide comprehensive analyses that reveal key limitations and opportunities in current approaches to grounded visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。