用自监督强化学习提升大模型空间理解能力,无需人工标注。
Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- 从普通图像自动构建五类空间预训练任务
- 在7个基准上分别提升3.89%和4.63%准确率
- 适合希望增强视觉模型空间推理的开发者
空间理解仍是大型视觉语言模型(LVLMs)的短板。现有监督微调(SFT)和近期基于可验证奖励的强化学习(RLVR)方法依赖昂贵标注、专用工具或受限环境,限制了规模。本文提出Spatial-SSRL,一种自监督强化学习范式,直接从普通RGB或RGB-D图像中提取可验证信号。该方法自动构建五类预训练任务,涵盖2D与3D空间结构:打乱图像块重排、翻转块识别、缺失块补全、区域深度排序及相对3D位置预测。这些任务提供易验证的真值答案,无需人工或LVLM标注。在图像与视频场景下的七个空间理解基准上,Spatial-SSRL相较于Qwen2.5-VL基线平均提升4.63%(3B模型)和3.89%(7B模型)准确率。结果表明,简单且内在的监督可实现大规模RLVR,为提升LVLM的空间智能提供可行路径。
原文摘要 · Abstract (English)
Spatial understanding remains a weakness of Large Vision-Language Models (LVLMs). Existing supervised fine-tuning (SFT) and recent reinforcement learning with verifiable rewards (RLVR) pipelines depend on costly supervision, specialized tools, or constrained environments that limit scale. We introduce Spatial-SSRL, a self-supervised RL paradigm that derives verifiable signals directly from ordinary RGB or RGB-D images. Spatial-SSRL automatically formulates five pretext tasks that capture 2D and 3D spatial structure: shuffled patch reordering, flipped patch recognition, cropped patch inpainting, regional depth ordering, and relative 3D position prediction. These tasks provide ground-truth answers that are easy to verify and require no human or LVLM annotation. Training on our tasks substantially improves spatial reasoning while preserving general visual capabilities. On seven spatial understanding benchmarks in both image and video settings, Spatial-SSRL delivers average accuracy gains of 4.63% (3B) and 3.89% (7B) over the Qwen2.5-VL baselines. Our results show that simple, intrinsic supervision enables RLVR at scale and provides a practical route to stronger spatial intelligence in LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。