arXiv:2503.14939cs.CV2025-03ICCV被引 12

评测多模态大模型的数字感知能力,发现普遍远低于人类水平。

VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

  • 构建涵盖7类视觉数感属性的1900题多选基准测试
  • 17个模型表现均显著低于人类,大模型仅小幅提升
  • 适合关注多模态推理与数学能力评估的研究者

多模态大模型能否具备类似人类的直观数字感知能力?为此,我们提出了视觉数字感知基准(VisNumBench),用于评估多模态大模型在多种视觉数字任务中的表现。该基准包含约1900个来自合成与真实图像的多选问答对,覆盖七种视觉数字属性和四类数字估计任务。实验发现:(i)我们测试的17个模型(包括Qwen2.5-VL、InternVL2.5等开源模型及GPT-4o、Gemini 2.0 Flash等闭源模型)在数字感知任务中表现显著低于人类水平;(ii)多模态数学模型与多模态思维链(CoT)模型未展现明显性能提升;(iii)参数量更大、通用能力更强的模型表现出微弱优势。我们认为VisNumBench将为社区提供重要资源,推动多模态大模型数字感知能力的发展。代码与数据集见https://wwwtttjjj.github.io/VisNumBench/。

原文摘要 · Abstract (English)

Can Multimodal Large Language Models (MLLMs) develop an intuitive number sense similar to humans? Targeting this problem, we introduce Visual Number Benchmark (VisNumBench) to evaluate the number sense abilities of MLLMs across a wide range of visual numerical tasks. VisNumBench consists of about 1,900 multiple-choice question-answer pairs derived from both synthetic and real-world visual data, covering seven visual numerical attributes and four types of visual numerical estimation tasks. Our experiments on VisNumBench led to the following key findings: (i) The 17 MLLMs we tested, including open-source models such as Qwen2.5-VL and InternVL2.5, as well as proprietary models like GPT-4o and Gemini 2.0 Flash, perform significantly below human levels in number sense-related tasks. (ii) Multimodal mathematical models and multimodal chain-of-thought (CoT) models did not exhibit significant improvements in number sense abilities. (iii) Stronger MLLMs with larger parameter sizes and broader general abilities demonstrate modest gains in number sense abilities. We believe VisNumBench will serve as a valuable resource for the research community, encouraging further advancements in enhancing MLLMs' number sense abilities. Code and dataset are available at https://wwwtttjjj.github.io/VisNumBench/.

多模态数字感知大模型评测视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。