用轻量提示词让视觉语言动作模型跨机器人高效学习
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

- 引入可学习嵌入作为各数据源的具身提示,实现跨平台统一建模
- 0.9B模型在6个仿真和3个真实机器人上均达当前最优性能
- 架构简洁可扩展,适合多机器人场景快速适配与部署
成功的通用视觉-语言-动作(VLA)模型依赖于在多样机器人平台上大规模、跨具身、异构数据集上的有效训练。为更好地利用丰富多样的机器人数据中的异质性,我们提出一种新型软提示方法,仅添加少量参数,将提示学习思想融入跨具身机器人学习,并为每种不同数据源引入独立的可学习嵌入。这些嵌入作为具身特异性提示,统一赋能VLA模型有效挖掘不同具身特征。我们的新模型X-VLA采用基于流匹配的简洁架构,完全依赖软提示化的标准Transformer编码器,兼具可扩展性与简洁性。在6个仿真环境及3个真实机器人上评估的0.9B版本——X-VLA-0.9B,在一系列基准测试中均达到最先进水平,展现出从灵活灵巧到跨具身、环境与任务的快速适应等广泛能力优势。
原文摘要 · Abstract (English)
Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse robotic data sources, we propose a novel Soft Prompt approach with minimally added parameters, by infusing prompt learning concepts into cross-embodiment robot learning and introducing separate sets of learnable embeddings for each distinct data source. These embeddings serve as embodiment-specific prompts, which in unity empower VLA models with effective exploitation of varying cross-embodiment features. Our new X-VLA, a neat flow-matching-based VLA architecture, relies exclusively on soft-prompted standard Transformer encoders, enjoying both scalability and simplicity. Evaluated across 6 simulations as well as 3 real-world robots, our 0.9B instantiation-X-VLA-0.9B simultaneously achieves SOTA performance over a sweep of benchmarks, demonstrating superior results on a wide axes of capabilities, from flexible dexterity to quick adaptation across embodiments, environments, and tasks. Website: https://thu-air-dream.github.io/X-VLA/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。