arXiv:2606.31645cs.CV2026-06

让AI在机器人任务中更准判断空间关系,靠两个推理激活技巧。

Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

论文配图:Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning
图 1 · 摘自论文原文
  • 用特殊指令触发深度思考,配合任务提示提升推理质量。
  • 解决相机视角与物体自身方向的混淆问题,准确率提升至80.9%。
  • 无需训练,适配复杂空间任务,适合机器人视觉研究者。

视觉语言模型虽具备强大感知能力,但在具身任务中常难以进行空间推理。我们提交RoboSpatialBrain至CVPR 2026的具身推理工作坊机器人空间挑战赛,基于RoboBrain2.5-8B-NV构建。该系统采用两种无需训练的推理阶段机制:强制使用<think>前缀激活策略,结合任务特化后提示,激发对上下文与兼容性任务的主动推理;以及显式参考系重定向流程,解决上下文任务中的相机中心与物体中心歧义。我们还探索了在兼容性数据上微调RoboBrain2.5的效果,并分析其与提示设计的交互关系。RoboSpatialBrain在RoboSpatial-Home数据集上取得80.9%的整体成功率,位居第一。代码已开源:https://github.com/YuxiangXie2003/RoboSpatialBrain。

原文摘要 · Abstract (English)

Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to the RoboSpatial Challenge at the Embodied Reasoning in Action Workshop, CVPR 2026, built on RoboBrain2.5-8B-NV. RoboSpatialBrain combines two training-free, inference-time mechanisms: a forced <think> prefix activation strategy paired with a task-specific post-prompt that elicits deliberate reasoning on context and compatibility tasks, and an explicit reference-frame redirection pipeline that resolves camera-centric and object-centric ambiguity for context tasks. We additionally explore fine-tuning RoboBrain2.5 on compatibility data and present a detailed analysis of its interaction with prompting. RoboSpatialBrain achieved first place in the RoboSpatial Challenge, with an overall success rate of 80.9\% on RoboSpatial-Home. Code is available at https://github.com/YuxiangXie2003/RoboSpatialBrain.

空间推理具身智能提示工程机器人视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。