让机器人精准理解复杂3D空间指令并动态推理位置。
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

- 用深度编码器增强视觉语言模型的3D空间感知能力。
- 通过强化学习实现最多5步的空间推理,准确率超89.6%。
- 适合需要复杂空间交互的机器人系统开发与研究。
空间指代是具身机器人与三维物理世界交互的基础能力。尽管预训练视觉语言模型(VLMs)强大,现有方法仍难以准确理解复杂3D场景并动态推理指令所指示的位置。为此,我们提出RoboRefer,一种3D感知的VLM,通过监督微调(SFT)集成解耦的专用深度编码器,首次实现精确的空间理解。此外,RoboRefer通过强化微调(RFT)实现泛化多步空间推理,采用针对空间指代任务设计的度量敏感过程奖励函数。为支持SFT和RFT训练,我们构建了包含2000万条问答对(是之前数据集的2倍)、涵盖31种空间关系(前为15种)并支持最多5步复杂推理的大型数据集RefSpatial。同时引入RefSpatial-Bench,填补多步推理评估空白。实验表明,经SFT训练的RoboRefer达到89.6%平均成功率;经RFT训练的版本显著超越所有基线,甚至在RefSpatial-Bench上比Gemini-2.5-Pro高出17.4%平均准确率。值得注意的是,RoboRefer可与多种控制策略集成,在杂乱真实场景中驱动多种机器人(如UR5、G1人形机器人)完成长时序动态任务。
原文摘要 · Abstract (English)
Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained vision language models (VLMs), recent approaches are still not qualified to accurately understand the complex 3D scenes and dynamically reason about the instruction-indicated locations for interaction. To this end, we propose RoboRefer, a 3D-aware VLM that can first achieve precise spatial understanding by integrating a disentangled but dedicated depth encoder via supervised fine-tuning (SFT). Moreover, RoboRefer advances generalized multi-step spatial reasoning via reinforcement fine-tuning (RFT), with metric-sensitive process reward functions tailored for spatial referring tasks. To support SFT and RFT training, we introduce RefSpatial, a large-scale dataset of 20M QA pairs (2x prior), covering 31 spatial relations (vs. 15 prior) and supporting complex reasoning processes (up to 5 steps). In addition, we introduce RefSpatial-Bench, a challenging benchmark filling the gap in evaluating spatial referring with multi-step reasoning. Experiments show that SFT-trained RoboRefer achieves state-of-the-art spatial understanding, with an average success rate of 89.6%. RFT-trained RoboRefer further outperforms all other baselines by a large margin, even surpassing Gemini-2.5-Pro by 17.4% in average accuracy on RefSpatial-Bench. Notably, RoboRefer can be integrated with various control policies to execute long-horizon, dynamic tasks across diverse robots (e,g., UR5, G1 humanoid) in cluttered real-world scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。