arXiv:2510.08531cs.CVcs.AI2025-10被引 61

通过分阶段训练提升视觉语言模型的空间推理能力

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models

  • 从物体定位逐步构建空间感知,再发展复杂推理能力
  • 在26,610样本数据集上训练,性能比基线高23.4%
  • 适合需要强空间理解能力的多模态应用开发者

空间推理仍是视觉语言模型(VLMs)的核心挑战,现有方法虽有进展,但表现仍不稳定。我们发现根本原因在于缺乏从感知到理解的层次基础。为此,提出一种渐进式训练方法:构建包含26,610个样本的SpatialLadder-26k多模态数据集,覆盖物体定位、单图、多视角及视频任务;设计三阶段训练框架——先通过物体定位建立空间感知,再通过多维空间任务发展空间理解,最后以可验证奖励的强化学习增强复杂推理。由此训练出的3B参数模型SpatialLadder,在空间推理基准上平均性能比基线提升23.4%,超越GPT-4o 20.8%、Gemini-2.0-Flash 10.1%;在域外测试集上仍保持7.2%的提升,证明从感知到推理的渐进训练对构建鲁棒空间智能至关重要。

原文摘要 · Abstract (English)

Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We identify that this limitation stems from a critical gap: existing methods attempt to learn spatial reasoning directly without establishing the hierarchical foundations of perception and understanding. To address this challenge, we present a comprehensive methodology for building spatial intelligence progressively. We introduce SpatialLadder-26k, a multimodal dataset containing 26,610 samples spanning object localization, single image, multi-view, and video spatial reasoning tasks, constructed through a standardized pipeline that ensures systematic coverage across modalities. Building on this dataset, we design a three-stage progressive training framework that (1) establishes spatial perception through object localization, (2) develops spatial understanding through multi-dimensional spatial tasks, and (3) strengthens complex reasoning via reinforcement learning with verifiable rewards. This approach yields SpatialLadder, a 3B-parameter model that achieves state-of-the-art performance on spatial reasoning benchmarks, with 23.4% average improvement over the base model, surpassing GPT-4o by 20.8% and Gemini-2.0-Flash by 10.1%. Notably, SpatialLadder maintains strong generalization with 7.2% improvement on out-of-domain benchmarks, demonstrating that progressive training from perception to reasoning is essential for robust spatial intelligence.

空间推理视觉语言模型渐进训练多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。