通过关键点对齐与分段强化学习,显著减少大模型空间推理中的错误累积。
STAR: Mitigating Cascading Errors in Spatial Reasoning via Turn-point Alignment and Segment-level DPO
- 基于拓扑锚点设计两阶段框架,先学习空间语义再优化自修正能力。
- 在RedMaze-23K数据集上,32B模型准确率达29.27%,接近GPT-4的82.4%。
- 适合研究空间推理、路径规划及大模型自我纠错机制的学者参考。
结构化空间导航是评估大语言模型空间推理能力的核心基准。现有方法如可视化思维(VoT)在复杂拓扑环境中易产生错误累积。为此,我们提出STAR框架,基于拓扑锚点设计两阶段流程,并构建了包含人类启发式转弯点标注的RedMaze-23K数据集。第一阶段通过监督微调帮助模型内化空间语义并剔除冗余路径;第二阶段采用空间感知的分段直接偏好优化(SDPO),提升长程导航中的自修正能力。实验表明,STAR在开源模型中达到领先性能:其32B版本准确率为29.27%,优于DeepSeek-V3(25.00%),达到GPT-4性能的82.4%。
原文摘要 · Abstract (English)
Structured spatial navigation is a core benchmark for Large Language Models (LLMs) spatial reasoning. Existing paradigms like Visualization-of-Thought (VoT) are prone to cascading errors in complex topologies. To solve this, we propose STAR, a two-stage framework grounded on topological anchors, and introduce the RedMaze-23K dataset with human-inspired turnpoint annotations. The first stage uses supervised fine-tuning to help models internalize spatial semantics and prune redundant paths. The second adopts Spatial-aware Segment-level Direct Preference Optimization (SDPO) to refine self-correction in long-horizon navigation. Experiments show STAR achieves state-of-the-art performance among open-source models: its 32B variant outperforms DeepSeek-V3 (29.27% vs. 25.00%) and reaches 82.4% of GPT-4's performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。