arXiv:2410.14702cs.AIcs.CL2024-10中稿 · NeurIPS被引 24

评测多模态模型在复杂图文推理中的表现,发现其空间理解能力严重不足。

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

  • 构建5000张高质图文混合题,覆盖10类认知任务
  • 顶尖模型最高仅41%正确率,普遍难解空间关系
  • 对比图文与纯文本提示,揭示模型不真正理解图像

多模态大语言模型在多个领域展现出强大的问题解决能力,但其视觉理解与抽象推理能力仍缺乏有效评估。为此,我们提出PolyMATH,一个旨在评估多模态大模型通用认知推理能力的挑战性基准。PolyMATH包含5000张手动收集的高质量图像,涵盖模式识别、空间推理、相对推理等10个不同类别。我们对15个主流多模态大模型进行了全面定量评估,采用四种不同的提示策略(包括Chain-of-Thought和Step-Back)。最佳模型成绩分别为:Claude-3.5 Sonnet约41%,GPT-4o约36%,Gemini-1.5 Pro约27%,凸显了题目在逻辑与视觉上的复杂性。细粒度错误分析表明,模型在理解空间关系和进行深层推理方面存在明显困难。消融实验进一步显示,当以文本描述替代图像时,模型性能仅提升约4%,说明模型并未真正理解图像中的空间信息,易产生逻辑错误。最后,我们评估了OpenAI o1系列模型,其表现仅达人类基线水平,验证了该基准的难度。PolyMATH的结果揭示了多模态推理仍有巨大提升空间,并为未来模型发展提供独特洞见。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs. PolyMATH comprises 5,000 manually collected high-quality images of cognitive textual and visual challenges across 10 distinct categories, including pattern recognition, spatial reasoning, and relative reasoning. We conducted a comprehensive, and quantitative evaluation of 15 MLLMs using four diverse prompting strategies, including Chain-of-Thought and Step-Back. The best scores achieved on PolyMATH are ~41%, ~36%, and ~27%, obtained by Claude-3.5 Sonnet, GPT-4o and Gemini-1.5 Pro respectively - highlighting the logical and visual complexity of these questions. A further fine-grained error analysis reveals that these models struggle to understand spatial relations and perform drawn-out, high-level reasoning. This is further strengthened by our ablation study estimating MLLM performance when given textual descriptions in place of diagrams. As evidenced by ~4% improvement over textual descriptions as opposed to actual images, we discover that models do not truly comprehend visual diagrams and the spatial information therein, and are thus prone to logical errors. Finally, we evaluate the OpenAI o1 models and find that their performance only matches the human baseline, highlighting the difficulty of the benchmark. The results on PolyMATH highlight the room for improvement in multi-modal reasoning and provide unique insights to guide the development of future MLLMs.

多模态推理图文理解认知评测模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。