测试模型空间路径推理能力,发现现有模型表现远低于人类。
SPaRC: A Spatial Pathfinding Reasoning Challenge
- 构建1000个二维网格路径谜题,需结合算术与几何规则逐步规划。
- 人类准确率达98.0%,而最强模型o4-mini仅15.8%,硬题下低至1.1%。
- 模型常生成无效路径,且无法随难度提升调整计算资源,适合研究者改进空间推理。
现有推理数据集趋于饱和,难以评估抽象、多步问题,尤其路径规划与复杂规则约束满足。我们提出SPaRC(空间路径推理挑战),包含1000个二维网格路径规划谜题,用于评估空间与符号推理能力,要求基于算术与几何规则进行逐步规划。人类在整体任务中达到98.0%准确率(困难题为94.5%),而最佳推理模型如o4-mini表现不佳,整体准确率仅15.8%,困难题低至1.1%。模型生成无效路径比例超过50%(以o4-mini为例),且其推理过程显示存在导航与空间逻辑错误。与人类不同,模型无法根据题目难度增加测试时计算量。允许模型多次尝试解题可提升准确率,表明通过改进训练和高效测试时扩展方法,有望提升空间推理能力。SPaRC可作为揭示模型空间推理局限性的窗口,推动面向抽象多步问题求解的新方法发展。
原文摘要 · Abstract (English)
Existing reasoning datasets saturate and fail to test abstract, multi-step problems, especially pathfinding and complex rule constraint satisfaction. We introduce SPaRC (Spatial Pathfinding Reasoning Challenge), a dataset of 1,000 2D grid pathfinding puzzles to evaluate spatial and symbolic reasoning, requiring step-by-step planning with arithmetic and geometric rules. Humans achieve near-perfect accuracy (98.0%; 94.5% on hard puzzles), while the best reasoning models, such as o4-mini, struggle (15.8%; 1.1% on hard puzzles). Models often generate invalid paths (>50% of puzzles for o4-mini), and reasoning tokens reveal they make errors in navigation and spatial logic. Unlike humans, who take longer on hard puzzles, models fail to scale test-time compute with difficulty. Allowing models to make multiple solution attempts improves accuracy, suggesting potential for better spatial reasoning with improved training and efficient test-time scaling methods. SPaRC can be used as a window into models' spatial reasoning limitations and drive research toward new methods that excel in abstract, multi-step problem-solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。