arXiv:2504.12597cs.CL2025-04被引 25

首个评估多模态模型几何推理能力的双语基准,聚焦原理识别与应用。

GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning

  • 构建五级几何原理框架,覆盖平面与立体几何,系统化评估推理机制。
  • 1789道标注题目中,Gemini-2.0-pro-flash表现最佳,整体得分65.3。
  • 揭示当前顶尖模型在原理识别与应用上仍存瓶颈,适合研究多模态推理者参考。

几何问题求解(GPS)是一项需要视觉理解与符号推理相结合的挑战性任务,能有效衡量多模态大语言模型(MLLMs)的推理能力。人类通过在视觉情境中准确识别并灵活运用几何原理展现出强大推理能力,但现有基准未能同时评估这一类人化的几何推理机制,成为评估模型处理GPS能力的关键缺口。为此,我们提出GeoSense,首个系统性评估MLLM几何推理能力的双语基准。GeoSense包含涵盖平面与立体几何的五级分层几何原理框架、1789道精心标注的问题数据集,以及创新的评估策略。在多种开源与闭源MLLM上进行的广泛实验表明,Gemini-2.0-pro-flash表现最优,整体得分为65.3。深入分析显示,几何原理的识别与应用仍是领先模型的瓶颈,共同制约其推理能力。这些发现凸显GeoSense在引导未来MLLM几何推理能力发展中的潜力,为实现更鲁棒、类人的智能推理铺平道路。

原文摘要 · Abstract (English)

Geometry problem-solving (GPS), a challenging task requiring both visual comprehension and symbolic reasoning, effectively measures the reasoning capabilities of multimodal large language models (MLLMs). Humans exhibit strong reasoning ability in this task through accurate identification and adaptive application of geometric principles within visual contexts. However, existing benchmarks fail to jointly assess both dimensions of the human-like geometric reasoning mechanism in MLLMs, remaining a critical gap in assessing their ability to tackle GPS. To this end, we introduce GeoSense, the first comprehensive bilingual benchmark designed to systematically evaluate the geometric reasoning abilities of MLLMs through the lens of geometric principles. GeoSense features a five-level hierarchical framework of geometric principles spanning plane and solid geometry, an intricately annotated dataset of 1,789 problems, and an innovative evaluation strategy. Through extensive experiments on GeoSense with various open-source and closed-source MLLMs, we observe that Gemini-2.0-pro-flash performs best, achieving an overall score of $65.3$. Our in-depth analysis reveals that the identification and application of geometric principles remain a bottleneck for leading MLLMs, jointly hindering their reasoning abilities. These findings underscore GeoSense's potential to guide future advancements in MLLMs' geometric reasoning capabilities, paving the way for more robust and human-like reasoning in artificial intelligence.

几何推理多模态模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。