提出可直接指导机器人动作的可视化空间定位接口,提升多模态决策与执行的准确性和透明度。
Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

- 通过几何专用头直接输出点位、物体功能热图和视觉轨迹,避免文本序列化空间信息
- 在桥接任务中实现72.9%成功率,无需特定微调,且控制时间减少20倍以上
- 适用于真实机器人部署,显著提升自主执行成功率至80.7%,适合需要高精度空间交互的场景
视觉-语言-动作(VLA)模型通常通过自回归文本坐标或模糊的动作标记实现空间定位,导致多模态推理与机器人执行之间接口脆弱。本文提出Pointing-VLA,一种基于Embodied-R1的类型化隐藏状态空间读出机制,采用几何专用头直接预测归一化点位、对象功能性定位(OFG)热图和视觉轨迹,无需将几何信息序列化为文本。在Bridge/WidowX及物理抓取放置部署中,明确执行契约将PICK绑定到源条件化的OFG,PLACE绑定到点位,提供阶段对齐的空间目标。Pointing-VLA在评估的四任务集上平均达到72.9%的成功率,未使用Bridge特化微调,且在支持碰撞检测的CuRobo执行下表现优异。点位与OFG在原数据集和跨数据集评估中展现出互补优势。OFG/接触读出可迁移至NORA-1.5,在保持或提升成功率的同时,控制器记录时间减少超过20倍;类型化头在共享外部套件上比Embodied-R1文本解码快6.68–6.90倍。当作为π₀.₅动作策略的空间引导时,其将真实机器人自主成功率从52.7%提升至80.7%,覆盖三个视觉场景。结果表明,类型化空间读出是高效、可解释的具身推理与机器人执行间接口。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. For the evaluated Bridge/WidowX and physical pick-place deployments, an explicit execution contract assigns PICK to source-conditioned OFG and PLACE to Pointing, providing direct stage-aligned spatial targets. Pointing-VLA achieves SOTA performance on Bridge/WidowX, averaging 72.9\% across the evaluated four-task set without Bridge-specific finetuning under collision-enabled CuRobo execution. Pointing and OFG show complementary strengths across native and cross-dataset evaluations. The OFG/contact readout transfers to NORA-1.5, preserving or improving success while reducing recorded controller time by more than 20$\times$; typed heads are also 6.68--6.90$\times$ faster than Embodied-R1 text decoding on a shared external suite. When integrated as spatial guidance for a $π_{0.5}$ action policy, Pointing-VLA raises autonomous real-robot success from 52.7\% to 80.7\% across three visual contexts. These results establish typed spatial readouts as an efficient, inspectable interface between embodied reasoning and robot execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。