提出新数据集与方法,提升大模型对几何关系的理解能力。
Do Large Language Models Truly Understand Geometric Structures?
- 设计GeomRel数据集,聚焦几何关系识别关键步骤。
- 多模型测试发现大模型在几何理解上存在显著局限。
- 提出GeoCoT方法,显著提升几何关系识别准确率。
几何能力是大型语言模型(LLMs)面临的重要挑战,因其需要高级空间理解与抽象思维。现有数据集主要评估模型的最终答案,无法真正衡量其对几何结构的理解,因为模型可能偶然得出正确答案。为填补这一空白,我们提出了GeomRel数据集,通过隔离问题求解中的核心步骤——几何关系识别,来评估LLMs对几何结构的真实理解。基于该基准,我们对多种LLMs进行了全面评估,揭示了其在几何理解上的关键缺陷。此外,我们提出了几何思维链(GeoCoT)方法,增强模型识别几何关系的能力,显著提升了性能。
原文摘要 · Abstract (English)
Geometric ability is a significant challenge for large language models (LLMs) due to the need for advanced spatial comprehension and abstract thinking. Existing datasets primarily evaluate LLMs on their final answers, but they cannot truly measure their true understanding of geometric structures, as LLMs can arrive at correct answers by coincidence. To fill this gap, we introduce the GeomRel dataset, designed to evaluate LLMs' understanding of geometric structures by isolating the core step of geometric relationship identification in problem-solving. Using this benchmark, we conduct thorough evaluations of diverse LLMs and identify key limitations in understanding geometric structures. We further propose the Geometry Chain-of-Thought (GeoCoT) method, which enhances LLMs' ability to identify geometric relationships, resulting in significant performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。