模型提前关注关键物体,提升跨场景泛化能力
Pelican-VLA 0.5: Attending Before Acting Benefits Generalization

- 通过可学习的瓶颈令牌,让视觉注意力聚焦操作相关区域
- 无需标注或微调,在未见场景和机器人上仍保持精准注意力
- 适合需要强泛化能力的机器人决策研究者
本文介绍Pelican-VLA 0.5,一个统一的视觉-语言-动作模型,整合了视觉语言理解、未来帧生成与动作预测。该模型实现注意力级泛化:无需物体标注、分割掩码、注意力监督或任务特定微调,其动作路径已能聚焦于操作相关的物体与接触区域。这种行为在未见场景和未见机器人形态中依然有效,显著优于其他开源VLA基线。我们验证该能力源于感知与动作间插入的可学习瓶颈令牌:通过紧凑瓶颈路由任务相关视觉信息,令牌接口在预训练中诱导出以操作为中心的注意力,并对不同策略结构(包括MoT风格架构)保持有效性。
原文摘要 · Abstract (English)
In this report, we present Pelican-VLA 0.5, a unified VLA model that integrates vision-language understanding, future-frame generation, and action prediction within a single architecture. Pelican-VLA 0.5 achieves attention-level generalization: without object annotations, segmentation masks, attention supervision, or task-specific fine-tuning, its action pathway already focuses on the manipulation-relevant object and contact region. This behavior persists across unseen scenes and unseen robot embodiments, and is substantially stronger than in other open-source VLA baselines. We verify that this ability originates from the learnable Bottleneck Token inserted between perception and action: by routing task-relevant visual information through a compact bottleneck, the tokens interface induces manipulation-centric attention during pre-training and remains effective across different policy structures, including a MoT-style architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。