让机器人通过视觉语言模型实现精准空间追踪与度量。
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics
- 构建3D感知的视觉语言模型,融合编码器与回归解码器增强尺度感知。
- 通过强化学习优化多步推理,平均成功率79.1%,在复杂场景中表现领先。
- 支持多种机器人执行长时序动态任务,适合真实环境下的智能控制应用。
空间追踪是机器人实现具身交互的基础能力,需结合多步度量推理、复杂空间指代和真实世界度量,但现有方法难以应对这一组合性挑战。为此,我们提出RoboTracer,一种3D感知的视觉语言模型,首次通过通用空间编码器与回归监督解码器,在监督微调中提升尺度感知能力,实现3D空间指代与测量。此外,通过带有度量敏感过程奖励的强化微调(RFT),引导关键中间感知线索,实现多步度量推理,准确生成空间轨迹。为支持训练,我们构建了包含3000万问答对的大规模数据集TraceSpatial,覆盖室内外及桌面场景,支持最多9步的复杂推理。我们还提出TraceSpatial-Bench作为挑战性基准,填补评估空白。实验表明,RoboTracer在空间理解、测量与指代上超越基线,平均成功率达79.1%,在TraceSpatial-Bench上性能显著优于Gemini-2.5-Pro,准确率高出36%。该模型可集成至多种控制策略,在复杂现实场景中驱动不同机器人(UR5、G1人形)完成长时序动态任务。
原文摘要 · Abstract (English)
Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spatial referring and real-world metric measurement. However, existing methods struggle with this compositional task. To this end, we propose RoboTracer, a 3D-aware VLM that first achieves both 3D spatial referring and measuring via a universal spatial encoder and a regression-supervised decoder to enhance scale awareness during supervised fine-tuning (SFT). Moreover, RoboTracer advances multi-step metric-grounded reasoning via reinforcement fine-tuning (RFT) with metric-sensitive process rewards, supervising key intermediate perceptual cues to accurately generate spatial traces. To support SFT and RFT training, we introduce TraceSpatial, a large-scale dataset of 30M QA pairs, spanning outdoor/indoor/tabletop scenes and supporting complex reasoning processes (up to 9 steps). We further present TraceSpatial-Bench, a challenging benchmark filling the gap to evaluate spatial tracing. Experimental results show that RoboTracer surpasses baselines in spatial understanding, measuring, and referring, with an average success rate of 79.1%, and also achieves SOTA performance on TraceSpatial-Bench by a large margin, exceeding Gemini-2.5-Pro by 36% accuracy. Notably, RoboTracer can be integrated with various control policies to execute long-horizon, dynamic tasks across diverse robots (UR5, G1 humanoid) in cluttered real-world scenes. Please see the project page at https://zhoues.github.io/RoboTracer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。