评测大模型空间推理中的视角理解能力,发现普遍短板并提出改进方法。
FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks
- 构建新基准FoREST,专门评估模型对空间视角的理解。
- 多类视角下大模型表现差异显著,影响图文生成准确性。
- 提出空间引导提示法,提升模型提取关键空间概念的能力。
空间推理是人类智能的核心能力之一。其中,空间参照系(Frame of Reference, FoR)用于识别空间描述的视角。尽管其重要性突出,但现有AI模型在该领域仍缺乏系统性研究与评估。为此,我们提出针对空间推理任务中参照系理解的评测基准FoREST,用于评估大语言模型(LLMs)在需要参照系理解的问题回答能力以及文本到图像模型的布局生成表现。实验结果揭示不同参照系类别间存在显著性能差距,影响了图文生成的准确性,暴露出模型在参照系理解上的关键缺陷。为改善这一问题,我们提出空间引导提示(Spatial-Guided Prompting),有效提升了模型提取关键空间概念的能力,显著改善了整体空间推理性能。
原文摘要 · Abstract (English)
Spatial reasoning is a fundamental aspect of human intelligence. One key concept in spatial cognition is the Frame of Reference, which identifies the perspective of spatial expressions. Despite its significance, FoR has received limited attention in AI models that need spatial intelligence. There is a lack of dedicated benchmarks and in-depth evaluation of large language models (LLMs) in this area. To address this issue, we introduce the Frame of Reference Evaluation in Spatial Reasoning Tasks (FoREST) benchmark, designed to assess FoR comprehension in LLMs. We evaluate LLMs on answering questions that require FoR comprehension and layout generation in text-to-image models using FoREST. Our results reveal a notable performance gap across different FoR classes in various LLMs, affecting their ability to generate accurate layouts for text-to-image generation. This highlights critical shortcomings in FoR comprehension. To improve FoR understanding, we propose Spatial-Guided prompting, which improves LLMs ability to extract essential spatial concepts. Our proposed method improves overall performance across spatial reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。