arXiv:2604.20834cs.RO2026-04被引 5

轻量级视觉语言动作模型,用世界知识提升机器人操作能力

PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

论文配图:PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
图 1 · 摘自论文原文
  • 两阶段训练:先预训练小型多模态模型,再注入动作相关语义与几何对齐
  • 在LIBERO-Plus上达成顶尖成功率,抗干扰能力强
  • 适合做机器人操作研究的开发者,代码数据已开源

近期视觉-语言-动作(VLA)模型推动了机器人操作的新发展,但现有方法效率有限且缺乏高层知识与空间意识。为此,我们提出PokeVLA,一种轻量级但强大的具身操作基础模型,能有效将视觉语言理解融入动作学习。框架采用两阶段训练:首先在包含240万样本的多模态数据集上预训练紧凑的视觉语言模型(PokeVLM),涵盖空间定位、可操作性及具身推理任务;其次通过多视角目标感知语义学习、几何对齐及新型动作专家,将操作相关表征注入动作空间。大量实验表明,PokeVLA在LIBERO-Plus基准和真实部署中均表现优异,成功率达领先水平,且在多种扰动下具备更强鲁棒性。为促进可复现性与社区进步,我们将开源代码、模型权重及预训练数据集构建脚本。项目页:https://getterupper.github.io/PokeVLA

原文摘要 · Abstract (English)

Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level knowledge and spatial awareness. To address these challenges, we propose PokeVLA, a lightweight yet powerful foundation model for embodied manipulation that effectively infuses vision-language understanding into action learning. Our framework introduces a two-stage training paradigm: first, we pre-train a compact vision-language model (PokeVLM) on a curated multimodal dataset of 2.4M samples encompassing spatial grounding, affordance, and embodied reasoning tasks; second, we inject manipulation-relevant representations into the action space through multi-view goal-aware semantics learning, geometry alignment, and a novel action expert. Extensive experiments demonstrate state-of-the-art performance on the LIBERO-Plus benchmark and in real-world deployment, outperforming comparable baselines in success rate and robustness under diverse perturbations. To foster reproducibility and community progress, we will open-source our code, model weights, and the scripts for the curated pre-training dataset. Project page: https://getterupper.github.io/PokeVLA

机器人操作视觉语言轻量模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。