arXiv:2411.19589cs.CL2024-11被引 12

测试大模型能否理解空间关系逻辑,发现其推理能力有限。

Can Large Language Models Reason about the Region Connection Calculus?

  • 用命名和无名空间关系测试大模型的推理能力
  • 30次重复实验显示模型表现不稳定且不准确
  • 适合关注AI空间认知局限的研究者阅读

定性空间推理是知识表示与推理领域的成熟方向,广泛应用于地理信息系统、机器人学和计算机视觉。近期大量声称大型语言模型(LLMs)具备推理能力。本文研究代表性LLMs在经典拓扑空间关系系统RCC-8中的定性空间推理能力。我们设计三组实验(重构组合表、对齐人类组合偏好、重构概念邻域),使用最先进LLMs进行测试;每组包含使用命名关系与匿名关系的对比实验,以评估模型是否依赖训练中获取的关系名称知识。所有实验重复30次,以衡量模型输出的随机性。

原文摘要 · Abstract (English)

Qualitative Spatial Reasoning is a well explored area of Knowledge Representation and Reasoning and has multiple applications ranging from Geographical Information Systems to Robotics and Computer Vision. Recently, many claims have been made for the reasoning capabilities of Large Language Models (LLMs). Here, we investigate the extent to which a set of representative LLMs can perform classical qualitative spatial reasoning tasks on the mereotopological Region Connection Calculus, RCC-8. We conduct three pairs of experiments (reconstruction of composition tables, alignment to human composition preferences, conceptual neighbourhood reconstruction) using state-of-the-art LLMs; in each pair one experiment uses eponymous relations and one, anonymous relations (to test the extent to which the LLM relies on knowledge about the relation names obtained during training). All instances are repeated 30 times to measure the stochasticity of the LLMs.

空间推理大模型知识表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。