arXiv:2412.21036cs.CL2024-12

评测大模型对几何形状与空间关系的理解能力,发现现有模型存在明显短板。

GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models

  • 构建专门针对几何感知的评测基准GePBench
  • 顶尖大模型在几何任务上准确率不足60%
  • 用该数据集训练可显著提升下游任务表现

几何形状在物理世界和人类认知中具有重要作用。尽管多模态大语言模型在视觉理解方面取得显著进展,但其对几何形状及其空间关系的识别能力——我们称之为几何感知——尚未得到系统性探索。为此,本文提出GePBench,一个专为评估多模态大模型几何感知能力而设计的新基准。大规模评测显示,即使最先进的模型在几何感知任务中也存在显著缺陷。此外,使用GePBench数据训练的模型在多种下游任务中表现出显著提升,凸显了几何感知在实现高级多模态应用中的关键作用。代码与数据集已开源至https://github.com/Changhao-Xiang/GePBench。

原文摘要 · Abstract (English)

Geometric shapes play important roles in both physical world and human cognition. While multimodal large language models (MLLMs) have made significant advancements in visual understanding, their abilities to recognize geometric shapes and their spatial relationships, which we term \emph{geometric perception}, are not explicitly and systematically explored. To address this gap, we introduce GePBench, a novel benchmark specifically designed to assess the geometric perception capabilities of MLLMs. Our extensive evaluations reveal that even the current state-of-the-art MLLMs exhibit significant deficiencies in geometric perception tasks. Furthermore, we show that models trained with GePBench data demonstrate considerable improvements on a wide range of downstream tasks, highlighting the critical role of geometric perception in enabling advanced multimodal applications. Our code and datasets are available at \href{https://github.com/Changhao-Xiang/GePBench}{https://github.com/Changhao-Xiang/GePBench}.

几何感知多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。