新基准评估多图推理能力,发现GPT-o1表现最佳且更稳定。
Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT
- 用多图任务+拒绝机制+熵值评估模型推理一致性
- GPT-o1准确率82.5%领先,QVQ-72B拒答准确率85.5%最高
- 小模型如Pixtral 12B在特定任务有潜力,Janus模型易受位置偏见影响
传统多模态大模型评估局限于单图推理,难以衡量上下文理解、推理稳定性与不确定性校准等关键能力。本文提出新基准,融合多图推理任务、基于拒绝的评估及位置偏差检测,并引入熵值作为量化推理一致性的新指标。评估涵盖Grok 3、ChatGPT-4o、ChatGPT-o1、Gemini 2.0 Flash Experimental、DeepSeek Janus系列、Qwen2.5-VL-72B-Instruct、QVQ-72B-Preview和Pixtral 12B共八种模型,在差异识别、图表解析等八个视觉推理任务上测试。结果表明,ChatGPT-o1综合准确率82.5%最高,拒绝准确率70.0%;Gemini 2.0 Flash Experimental以70.8%紧随其后。QVQ-72B-Preview拒答准确率达85.5%,表现优异。Pixtral 12B在特定任务中展现潜力(51.7%),而Janus模型因位置偏差与不确定性校准不足,呈现低拒绝准确率与高熵值(如Janus 7B:0.8392,Janus 1B:0.787),显示推理不稳定。研究还发现模型规模并非决定性因素,例如Grok 3虽参数量大但表现欠佳。该基准通过多图上下文、拒绝机制与熵度量,为下一代AI系统评估建立新标准。
原文摘要 · Abstract (English)
Traditional evaluations of multimodal large language models (LLMs) have been limited by their focus on single-image reasoning, failing to assess crucial aspects like contextual understanding, reasoning stability, and uncertainty calibration. This study addresses these limitations by introducing a novel benchmark that integrates multi-image reasoning tasks with rejection-based evaluation and positional bias detection. To evaluate these dimensions, we further introduce entropy as a novel metric for quantifying reasoning consistency across reordered answer variants. We applied this benchmark to assess Grok 3, ChatGPT-4o, ChatGPT-o1, Gemini 2.0 Flash Experimental, DeepSeek Janus models, Qwen2.5-VL-72B-Instruct, QVQ-72B-Preview, and Pixtral 12B across eight visual reasoning tasks, including difference spotting and diagram interpretation. Our findings reveal ChatGPT-o1 leading in overall accuracy (82.5\%) and rejection accuracy (70.0\%), closely followed by Gemini 2.0 Flash Experimental (70.8\%). QVQ-72B-Preview demonstrated superior rejection accuracy (85.5\%). Notably, Pixtral 12B (51.7\%) showed promise in specific domains, while Janus models exhibited challenges in bias and uncertainty calibration, reflected in low rejection accuracies and high entropy scores. High entropy scores in Janus models (Janus 7B: 0.8392, Janus 1B: 0.787) underscore their susceptibility to positional bias and unstable reasoning, contrasting with the low entropy and robust reasoning of ChatGPT models. The study further demonstrates that model size is not the sole determinant of performance, as evidenced by Grok 3 underperformance despite its substantial parameter count. By employing multi-image contexts, rejection mechanisms, and entropy-based consistency metrics, this benchmark sets a new standard for evaluating multimodal LLMs, enabling a more robust and reliable assessment of next-generation AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。