arXiv:2605.22183cs.ROcs.AI2026-05被引 1

将视觉原语引入机器人操作,提升任务成功率与泛化能力

Action with Visual Primitives

论文配图:Action with Visual Primitives
图 1 · 摘自论文原文
  • 用视觉原语作为中间表示,分离感知与动作决策
  • 实测成功率比pi_0.5高37.04%,数据效率更优
  • 适合需要空间组合泛化和物体迁移的机器人任务

视觉-语言-动作(VLA)模型已成为通用机器人操作的有前景范式。当前架构通常在单次前向传播中将语言指令与视觉观测映射为动作,但这种设计将指令理解、场景空间认知与运动控制耦合在单一学习目标中,导致动作专家需重复学习预训练视觉语言模型已具备的认知与感知能力,限制了学习效率与泛化性能。我们提出AVP(Action with Visual Primitives),一种端到端架构,采用以视觉原语为中心的接口:视觉语言模型推断下一阶段目标并生成视觉原语标记,由流匹配动作专家根据末端执行器运动学进行监督。真实机器人上针对通用抓取与放置任务的实验表明,AVP相较于pi_0.5成功率提升37.04%,并在数据效率、空间组合泛化与物体级迁移方面持续领先。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current architectures maps language instructions and visual observations to actions in a single forward pass. While conceptually simple, this formulation entangles instruction comprehension, spatial scene understanding, and motor control within a single learning objective. As a result, the action expert must implicitly relearn cognitive and perceptual capabilities already present in the pretrained VLM, which can limit both learning efficiency and generalization. We introduce AVP (Action with Visual Primitives), an end-to-end architecture that implements this visual-primitive-centric interface: the VLM infers the next-stage target and emits visual-primitive tokens that condition a flow-matching action expert, with supervision derived from end-effector kinematics. Real-robot experiments on general pick-and-place tasks show that AVP improves the success rate by 37.04% over pi_0.5 and outperforms other recent methods, with consistent gains in data efficiency, spatial-compositional generalization, and object-level transfer.

机器人操作视觉原语端到端泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。