arXiv:2608.07353cs.CLcs.AI2026-08

测试大模型对空间概念的抽象、组合与定位能力,发现其理解存在明显短板。

Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding

论文配图:Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
图 1 · 摘自论文原文
  • 设计针对空间概念的三类测试:抽象、组合、定位
  • 多模型实验显示当前大模型在概念组合上表现不佳
  • 适合研究模型认知机制或知识管理的学者参考

理解概念是实现泛化的核心。尽管大语言模型在多种任务上表现优异,但在真正理解概念方面仍存在不足。以往研究多依赖自然语言基准或限定范围的合成任务,这些方法常混淆多种能力,且对概念及其属性控制不精确。为实现对概念的可控探测,我们设计了针对抽象性、组合性和可定位性的测试。构建以方向、距离、拓扑等空间概念为核心的基准,采用问答任务作为代理评估方式。在多个大模型架构和训练策略下开展广泛实验,分析模型规模与结构对概念理解的影响。结果揭示当前大模型存在明显局限,并为提升其获取与组合结构化概念的能力提供了洞见。研究结果有助于重新设计基于概念的大模型,以增强信息检索与知识管理能力。代码将开源于 https://github.com/rd20karim/concept-probing。

原文摘要 · Abstract (English)

Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Models (LLMs) still struggle with genuine concept understanding. Prior work has evaluated conceptual understanding in LLMs using natural-language benchmarks or narrowly scoped synthetic tasks, but these settings often conflate multiple skills or lack precise control over the underlying concepts and their properties. To support controlled probing of concepts in LLMs, we design tests on their core properties: abstraction, compositionality, and groundness. We set up a concept-centric benchmark, targeting spatial concepts such as direction, distance, topology, and their compositions, and use question answering tasks serving as a proxy. We conduct extensive experiments across multiple LLM architectures and training regimes to analyze how model scale and design impact conceptual understanding. The results reveal clear limitations in current LLMs and provide insights into the factors shaping their ability to acquire and compose structured concepts. Our findings shed light on how concept-based LLMs can be redesigned for improved information access and knowledge management. The code will be available at https://github.com/rd20karim/concept-probing.

概念理解空间推理模型评测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。