arXiv:2509.18282cs.ROcs.AI2025-09被引 11

用视觉语言模型生成最小关键点表示,提升机器人抓取政策零样本泛化能力。

PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies

  • 利用视觉语言模型生成统一的点状中间表示,分离关注点与动作执行
  • 在20多个数据集上训练,3D策略真实场景性能提升41.4倍
  • 适用于大模型和小模型,适合需要强泛化的机器人控制任务

机器人操作策略常因需同时学习关注位置、执行动作及如何操作而难以泛化。本文提出PEEK(Policy-agnostic Extraction of Essential Keypoints),通过微调视觉语言模型(VLMs)生成统一的基于点的中间表示:1)末端执行器路径,明确应采取的动作;2)任务相关掩码,指示关注区域。这些标注直接叠加于机器人观测之上,使表示具有政策无关性,可跨架构迁移。为实现可扩展训练,构建自动标注流水线,在20+个机器人数据集(涵盖9种机器人形态)上生成标注数据。真实世界评估显示,PEEK持续提升零样本泛化能力,包括仅在仿真中训练的3D策略在真实世界性能提升41.4倍,大型视觉语言模型和小型操作策略均获得2–3.5倍收益。通过让VLMs承担语义与视觉复杂度,PEEK为操作策略提供最小但必需的线索——何处、何物、如何行动。

原文摘要 · Abstract (English)

Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-level reasoning about where and what can be offloaded to vision-language models (VLMs), leaving policies to specialize in how to act. We present PEEK (Policy-agnostic Extraction of Essential Keypoints), which fine-tunes VLMs to predict a unified point-based intermediate representation: 1. end-effector paths specifying what actions to take, and 2. task-relevant masks indicating where to focus. These annotations are directly overlaid onto robot observations, making the representation policy-agnostic and transferable across architectures. To enable scalable training, we introduce an automatic annotation pipeline, generating labeled data across 20+ robot datasets spanning 9 embodiments. In real-world evaluations, PEEK consistently boosts zero-shot generalization, including a 41.4x real-world improvement for a 3D policy trained only in simulation, and 2-3.5x gains for both large VLAs and small manipulation policies. By letting VLMs absorb semantic and visual complexity, PEEK equips manipulation policies with the minimal cues they need--where, what, and how. Website at https://peek-robot.github.io/.

机器人操控零样本泛化视觉语言模型关键点表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。