简单基线模型在多任务机器人任务中表现优异,证明复杂设计非必要。
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems

- 用极简架构控制变量,系统性测试视觉-语言-动作模型设计选择。
- 统一训练下跨多个基准表现强,真实场景任务超越π₀.₅达20%。
- 适合想快速验证新想法或避免过度工程化的研究者使用。
视觉-语言-动作(VLA)模型已成为构建通用机器人智能体的有前景范式。然而,现有方法在架构、训练数据、机器人配置和基准工程上差异巨大,导致研究碎片化。本文提出星型基线模型StarVLA-α,通过最小化架构与流程复杂度,减少实验混杂因素,实现可控分析。我们重新评估了动作建模策略、机器人特化预训练和接口工程等关键设计维度。在LIBERO、SimplerEnv、RoboTwin和RoboCasa四个基准上的统一多任务训练中,该简单基线仍保持高竞争力,表明强大视觉语言模型骨干搭配极少设计即可达成优异性能,无需依赖额外复杂结构或工程技巧。值得注意的是,单个通用模型在公开真实世界RoboChallenge基准上比π₀.₅提升20%。我们期望StarVLA-α成为未来VLA研究的可靠起点。代码将发布于https://github.com/starVLA/starVLA。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for building general-purpose robotic agents. However, the VLA landscape remains highly fragmented and complex: as existing approaches vary substantially in architectures, training data, embodiment configurations, and benchmark-specific engineering. In this work, we introduce StarVLA-$α$, a simple yet strong baseline designed to study VLA design choices under controlled conditions. StarVLA-$α$ deliberately minimizes architectural and pipeline complexity to reduce experimental confounders and enable systematic analysis. Specifically, we re-evaluate several key design axes, including action modeling strategies, robot-specific pretraining, and interface engineering. Across unified multi-benchmark training on LIBERO, SimplerEnv, RoboTwin, and RoboCasa, the same simple baseline remains highly competitive, indicating that a strong VLM backbone combined with minimal design is already sufficient to achieve strong performance without relying on additional architectural complexity or engineering tricks. Notably, our single generalist model outperforms $π_{0.5}$ by 20\% on the public real-world RoboChallenge benchmark. We expect StarVLA-$α$ to serve as a solid starting point for future research in the VLA regime. Code will be released at https://github.com/starVLA/starVLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。