分阶段控制机械臂:先定位后操作,提升复杂环境下的抓取成功率。
GloVLA: Let Geometry Move and Local VLA Interact for Robust Object-Centric Manipulation in Unstructured Environments

- 将动作分为几何定位与局部交互两阶段,分离任务难点
- 在真实机器人上成功率从35.6%提升至90.0%,推理时间减半
- 兼容多种模型,无需额外训练,适合工业级部署
视觉-语言-动作(VLA)模型在语言驱动的机器人操作中展现出良好泛化能力,但在非结构化环境中部署仍具挑战。单一端到端的VLA策略需同时处理长距离末端执行器移动和到达后的短时接触式交互,效率低且易失效。微小视觉偏移、干扰物、杂乱环境、遮挡或初始夹爪姿态不利,都会使策略脱离训练时的局部状态分布,导致任务失败。本文提出GloVLA,一种混合框架,将以物体为中心的操作明确分为两个互补阶段:几何运输控制器负责将末端执行器移动至交互预设区域,局部VLA策略仅处理短时交互阶段。GloVLA具有模型无关性,可无缝集成不同VLA主干网络,无需额外演示,不改变动作空间或成功判定标准。在标准LIBERO和LIBERO-Plus Object任务及新提出的LIBERO-Challenge基准(含杂乱、干扰物、光照变化、视觉偏移与遮挡)上的实验表明,相比全轨迹端到端方法,GloVLA显著提升任务成功率并大幅降低推理成本。在模拟环境中,全轨迹GR00T N1.6执行成功率降至20.9%,而GloVLA保持88.5%;在物理UR10e机器人上,总体成功率从35.6%提升至90.0%,平均推理时间超过一半。视频与更多结果见https://glovla-project.github.io/
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have shown promising generalization for language-conditioned robot manipulation, but deploying them in unstructured environments remains challenging. A single end-to-end VLA policy must simultaneously solve long-range transport of the end effector to task-relevant regions and short-horizon, contact-rich interaction upon arrival. This formulation is inefficient and brittle: small visual shifts, distractors, clutter, occlusions, or unfavorable initial gripper poses can push the policy outside the local state distribution in which it was trained, leading to task failure. We introduce GloVLA, a hybrid framework that explicitly separates object-centric manipulation into two complementary regimes: a geometric transport controller moves the end-effector into interaction-centric handoff regions, and local VLA policies handle only the short-horizon interaction phases. GloVLA is model-agnostic and can be integrated with different VLA backbones with no additional demonstrations and no changes to the action space or success predicate. Experiments on standard LIBERO and LIBERO-Plus Object tasks together with a newly introduced LIBERO-Challenge benchmark ettings with clutter, distractors,illumination changes, visual shifts, and obstruction show that GloVLA improves task success and substantially lowers VLA inference cost compared with full end-to-Challenge, full-trajectory GR00T N1.6execution degrades to 20.9% average success while GloVLA retains 88.5%; on a physical UR10e, overall success improves from 35.6% to 90.0% while mean inference time is more than halved. Videos and additional results are available at https://glovla-project.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。