测试大模型对几何形状的直观理解能力,发现顶尖模型准确率仅65%。
NoReGeo: Non-Reasoning Geometry Benchmark
- 设计2500个无需推理的几何题,纯靠空间感知解题
- 最先进模型在二分类任务中最高仅达65%准确率
- 表明当前模型缺乏原生几何认知,需专门训练
我们提出NoReGeo,一个新型基准,用于评估大语言模型(LLMs)在不依赖推理或代数计算的前提下对几何概念的内在理解能力。与现有主要考察代数推理能力的几何基准不同,NoReGeo关注模型是否能直接编码空间关系并识别几何属性。该基准包含2500个简单的几何问题,覆盖25个类别,每个问题均可在已知物体位置前提下仅通过本征几何理解求解。我们在包括GPT-4在内的多种前沿模型上进行了评估,发现即使最先进的系统在二分类任务中最高准确率也仅为65%。进一步的消融实验表明,这种几何理解能力无法仅通过微调获得,说明有效提升几何认知需要从初始阶段就采用专门训练方法。研究结果揭示了当前大模型在原生几何认知上的显著差距,为未来具备真正几何认知能力的模型研究奠定了基础。
原文摘要 · Abstract (English)
We present NoReGeo, a novel benchmark designed to evaluate the intrinsic geometric understanding of large language models (LLMs) without relying on reasoning or algebraic computation. Unlike existing benchmarks that primarily assess models' proficiency in reasoning-based geometry-where solutions are derived using algebraic methods-NoReGeo focuses on evaluating whether LLMs can inherently encode spatial relationships and recognize geometric properties directly. Our benchmark comprises 2,500 trivial geometric problems spanning 25 categories, each carefully crafted to be solvable purely through native geometric understanding, assuming known object locations. We assess a range of state-of-the-art models on NoReGeo, including frontier models like GPT-4, observing that even the most advanced systems achieve an overall maximum of 65% accuracy in binary classification tasks. Further, our ablation experiments demonstrate that such geometric understanding does not emerge through fine-tuning alone, indicating that effective training for geometric comprehension requires a specialized approach from the outset. Our findings highlight a significant gap in current LLMs' ability to natively grasp geometric concepts, providing a foundation for future research toward models with true geometric cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。