评测大模型视觉细粒度感知与因果推理能力,发现表现仍有限。
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
- 构建双难度多模态基准Argus Inspection,聚焦细节识别与常识推理。
- 26个主流模型在细粒度视觉任务中最高仅达46%准确率。
- 提出新评估框架,适合研究视觉认知与智能系统可信性的人参考。
随着多模态大语言模型(MLLMs)不断发展,其认知与推理能力取得显著进展,但在视觉细粒度感知和常识因果推理方面仍存在挑战。本文提出Argus Inspection多模态基准,包含两个难度层级,强调细致视觉识别,并融合真实世界常识理解以评估因果推理能力。基于此,我们设计了Eye of Panoptes评估框架,结合二元参数Sigmoid度量与指示函数,实现对意见类推理任务中模型响应的更全面评估。在26个主流MLLM上的实验表明,视觉细粒度推理的最高性能仅为0.46,显示出巨大提升空间。本研究为MLLM的持续优化提供了重要视角。
原文摘要 · Abstract (English)
As Multimodal Large Language Models (MLLMs) continue to evolve, their cognitive and reasoning capabilities have seen remarkable progress. However, challenges in visual fine-grained perception and commonsense causal inference persist. This paper introduces Argus Inspection, a multimodal benchmark with two levels of difficulty, emphasizing detailed visual recognition while incorporating real-world commonsense understanding to evaluate causal reasoning abilities. Expanding on it, we present the Eye of Panoptes framework, which integrates a binary parametric Sigmoid metric with an indicator function, enabling a more holistic evaluation of MLLMs' responses in opinion-based reasoning tasks. Experiments conducted on 26 mainstream MLLMs reveal that the highest performance in visual fine-grained reasoning reaches only 0.46, highlighting considerable potential for enhancement. Our research offers valuable perspectives for the continued refinement of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。