测试28个大模型对方向推理的能力,发现新模型仍不靠谱。
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
- 用模板生成测试题,覆盖不同视角和移动方式
- 多数模型在复杂场景下方向判断准确率不足
- 适合评估模型空间推理能力的研究者参考
我们研究了28个大型语言模型(LLMs)在给定特定情境下对基本方向(Cardinal Directions, CDs)进行推理的能力,使用基于模板构建的基准测试集进行广泛测试。模板支持多种变量,包括参与者的移动方式以及第一、第二或第三人称视角。结果显示,即使是较新的大型推理模型,在所有问题上也未能可靠地确定正确方向。本文总结并扩展了此前在COSIT-24上发表的工作。
原文摘要 · Abstract (English)
We investigate the abilities of 28 Large language Models (LLMs) to reason about cardinal directions (CDs) using a benchmark generated from a set of templates, extensively testing an LLM's ability to determine the correct CD given a particular scenario. The templates allow for a number of degrees of variation such as means of locomotion of the agent involved, and whether set in the first, second or third person. Even the newer Large Reasoning Models are unable to reliably determine the correct CD for all questions. This paper summarises and extends earlier work presented at COSIT-24.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。