绿机器人通用智能体,分阶段训练实现跨平台精准控制。
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
- 五阶段渐进式训练:从基础模型到具体机器人适配
- 3000小时示范数据+时序对齐,提升动作一致性
- 单策略控制多类机器人,适合工业与服务场景
我们提出Green-VLA,一种用于Green人形机器人真实部署的分阶段视觉-语言-动作(VLA)框架,兼顾多样化机器人形态的泛化能力。该框架采用五阶段课程:(L0) 基础视觉语言模型,(L1) 多模态定位,(R0) 多形态预训练,(R1) 形态特异性适配,(R2) 强化学习策略对齐。结合可扩展的数据处理流程(3000小时演示数据)、时序对齐与质量过滤,并使用统一的、具身感知的动作接口,使单一策略可控制人形机器人、移动操作臂和固定基座机械臂。推理时,通过任务进展预测、分布外检测及联合预测引导提升安全性和目标选择精度。在Simpler BRIDGE WidowX和CALVIN ABC-D上的实验,以及真实机器人测试均显示,强化学习对齐显著提升了成功率、鲁棒性与长程任务效率。
原文摘要 · Abstract (English)
We introduce Green-VLA, a staged Vision-Language-Action (VLA) framework for real-world deployment on the Green humanoid robot while maintaining generalization across diverse embodiments. Green-VLA follows a five stage curriculum: (L0) foundational VLMs, (L1) multimodal grounding, (R0) multi-embodiment pretraining, (R1) embodiment-specific adaptation, and (R2) reinforcement-learning (RL) policy alignment. We couple a scalable data-processing pipeline (3,000 hours of demonstrations) with temporal alignment and quality filtering, and use a unified, embodiment-aware action interface enabling a single policy to control humanoids, mobile manipulators, and fixed-base arms. At inference, the VLA controller is enhanced with episode-progress prediction, out-of-distribution detection, and joint-prediction-based guidance to improve safety and precise target selection. Experiments on Simpler BRIDGE WidowX and CALVIN ABC-D, as well as real-robot evaluations, show strong generalization and performance gains from RL alignment in success rate, robustness, and long-horizon efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。