测试多模态大模型解数学题能力,发现它们看图解题仍不成熟。
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
- 构建跨语言数学图文评测集,覆盖几何、代数等五类题型
- 最佳模型仅达人类水平的60%,图像信息利用率普遍偏低
- 谷歌Gemini和GPT-4o推理更结构化,其他模型易靠猜答
多模态大语言模型(MLLMs)虽具备视觉语言能力,但在处理图文数学问题方面仍缺乏系统评估。本文基于袋鼠数学竞赛风格,构建涵盖英语、法语、西班牙语和加泰罗尼亚语的多语言评测基准,评估GPT-4o、Pixtral、Qwen VL、Llama 3.2 Vision及Gemini 2.0 Flash等模型在几何、视觉代数、逻辑、模式与组合推理任务中的表现。实验发现:整体准确率中等,无模型在所有题型上均领先;多数模型在无图题上提升有限,部分性能几乎不变,表明对图示信息利用不足;不同语言与难度下差异显著,高级几何与组合题普遍困难;其中Gemini 2.0 Flash在有图任务中精度最高,其次为Qwen VL 2.5 72B与GPT-4o,但均未达人类水平。进一步分析显示,Gemini与GPT-4o展现出更稳定的结构化推理,而Pixtral与Llama常依赖启发式或随机猜测。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) promise advanced vision language capabilities, yet their effectiveness in visually presented mathematics remains underexplored. This paper analyzes the development and evaluation of MLLMs for mathematical problem solving, focusing on diagrams, multilingual text, and symbolic notation. We then assess several models, including GPT 4o, Pixtral, Qwen VL, Llama 3.2 Vision variants, and Gemini 2.0 Flash in a multilingual Kangaroo style benchmark spanning English, French, Spanish, and Catalan. Our experiments reveal four key findings. First, overall precision remains moderate across geometry, visual algebra, logic, patterns, and combinatorics: no single model excels in every topic. Second, while most models see improved accuracy with questions that do not have images, the gain is often limited; performance for some remains nearly unchanged without visual input, indicating underutilization of diagrammatic information. Third, substantial variation exists across languages and difficulty levels: models frequently handle easier items but struggle with advanced geometry and combinatorial reasoning. Notably, Gemini 2.0 Flash achieves the highest precision on image based tasks, followed by Qwen VL 2.5 72B and GPT 4o, though none approach human level performance. Fourth, a complementary analysis aimed at distinguishing whether models reason or simply recite reveals that Gemini and GPT 4o stand out for their structured reasoning and consistent accuracy. In contrast, Pixtral and Llama exhibit less consistent reasoning, often defaulting to heuristics or randomness when unable to align their outputs with the given answer options.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。