LingBot-VLA 2.0通过三方面改进,让机器人在真实场景中更智能地执行复杂任务。
From Foundation to Application: Improving VLA Models in Practice

- 用6万小时新数据重训,覆盖20种机器人配置和人类视角视频
- 支持头部、腰部、移动基座等多自由度,提升复杂任务处理能力
- 引入未来预测机制,增强时间推理,适合实际应用部署
尽管视觉-语言-动作(VLA)基础模型取得进展,但实验室与真实应用场景之间的差距仍阻碍其落地。为此,我们提出LingBot-VLA 2.0,从三个方向实现升级:(1) 任务与机体泛化能力。重构数据处理流程,新增约6万小时预训练数据,包括5万小时跨20种机器人配置的轨迹数据和1万小时第一人称人类视频;(2) 扩展动作空间,支持头部、腰部、移动基座及灵巧手的多自由度控制,使机器人可在实际场景中完成更复杂任务;(3) 引入预测动力学建模以增强时序推理,将未来预测作为代理任务,结合视频表征模型提供语义先验、深度估计模型提供几何线索。在GM-100基准上的通用设置评估验证了这些改进的有效性。得益于覆盖全身自由度的大规模预训练数据,LingBot-VLA 2.0在两种机器人平台上均展现出出色的跨机体长时程移动操作能力。
原文摘要 · Abstract (English)
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。