arXiv:2601.18692cs.ROcs.CV2026-01被引 83

基于真实数据训练的视觉语言动作模型,通用性强且训练高效。

A Pragmatic VLA Foundation Model

  • 用9种双臂机器人真实数据训练,覆盖2万小时场景
  • 在4个平台上完成100任务,每任务130次微调后表现领先
  • 代码库每秒处理261样本,速度提升1.5~2.8倍,适合部署

具备在机器人操作中广泛应用潜力的视觉-语言-动作(VLA)基础模型,需在任务与平台间实现良好泛化,并兼顾成本效率(如数据和GPU时长)。为此,我们构建了LingBot-VLA模型,采用来自9种常见双臂机器人配置的约2万小时真实世界数据。通过在4个机器人平台上进行系统评估,每个平台完成100项任务,每项任务进行130次微调后评估,该模型显著优于现有方法,展现出优异性能与广泛泛化能力。我们还开发了一个高效代码库,在8张GPU上达到每秒261样本的吞吐量,相比现有VLA代码库提速1.5~2.8倍(取决于所依赖的视觉语言模型基座)。上述特性使该模型非常适合实际部署。为推动机器人学习发展,我们开放提供代码、基础模型及基准数据,致力于支持更复杂任务并建立科学评估标准。

原文摘要 · Abstract (English)

Offering great potential in robotic manipulation, a capable Vision-Language-Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensuring cost efficiency (e.g., data and GPU hours required for adaptation). To this end, we develop LingBot-VLA with around 20,000 hours of real-world data from 9 popular dual-arm robot configurations. Through a systematic assessment on 4 robotic platforms, each completing 100 tasks with 130 post-training episodes per task, our model achieves clear superiority over competitors, showcasing its strong performance and broad generalizability. We have also built an efficient codebase, which delivers a throughput of 261 samples per second with an 8-GPU training setup, representing a 1.5~2.8$\times$ (depending on the relied VLM base model) speedup over existing VLA-oriented codebases. The above features ensure that our model is well-suited for real-world deployment. To advance the field of robot learning, we provide open access to the code, base model, and benchmark data, with a focus on enabling more challenging tasks and promoting sound evaluation standards.

机器人学习视觉语言动作基础模型高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。