arXiv:2503.00043cs.CVcs.AI2025-03ICLR被引 3

评测大模型跨图像抽象推理能力,发现当前模型表现远低于人类。

VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning

  • 设计动态开放题型,要求模型完成图像类比生成任务。
  • 最先进模型在复杂场景下准确率仅13%,简单任务也仅29%。
  • 揭示大模型缺乏高层次关系推理能力,适合关注模型局限的研究者。

多模态大语言模型(MLLMs)已成为融合视觉与文本信息的强大工具。尽管其在视觉理解基准上表现优异,但评估其跨多图的抽象推理能力仍具挑战。为此,我们提出VOILA——一个大规模、开放式、动态的基准,用于评估MLLMs的感知理解与抽象类比推理能力。该基准采用视觉领域的类比映射方法,要求模型在无预定义选项的情况下,生成一张能补全两组图像间类比关系的图片。实验表明,VOILA中的类比推理任务对现有MLLMs构成显著挑战。通过多步分析发现,当前模型难以理解图像间关系,高阶关系推理能力有限。值得注意的是,采用从简到繁的提示策略可提升性能。对开源模型与GPT-4o的全面评估显示,在文本回答任务中,最复杂场景的最佳准确率为13%(LLaMa 3.2),简单任务也仅为29%(GPT-4o),而人类在两类任务上的准确率均达70%。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To address this, we introduce VOILA, a large-scale, open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding and abstract relational reasoning. VOILA employs an analogical mapping approach in the visual domain, requiring models to generate an image that completes an analogy between two given image pairs, reference and application, without relying on predefined choices. Our experiments demonstrate that the analogical reasoning tasks in VOILA present a challenge to MLLMs. Through multi-step analysis, we reveal that current MLLMs struggle to comprehend inter-image relationships and exhibit limited capabilities in high-level relational reasoning. Notably, we observe that performance improves when following a multi-step strategy of least-to-most prompting. Comprehensive evaluations on open-source models and GPT-4o show that on text-based answers, the best accuracy for challenging scenarios is 13% (LLaMa 3.2) and even for simpler tasks is only 29% (GPT-4o), while human performance is significantly higher at 70% across both difficulty levels.

多模态类比推理模型评测视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。