arXiv:2602.18224cs.ROcs.LG2026-02被引 6

轻量级视觉语言动作模型,简单有效,性能超越大模型。

SimVLA: A Simple VLA Baseline for Robotic Manipulation

  • 分离感知与控制,用标准骨干+轻量动作头,设计简洁透明。
  • 仅0.5亿参数,在仿真任务上超过数十亿参数的模型。
  • 适合想对比新架构的开发者,可复现且能清晰归因性能提升。

视觉-语言-动作(VLA)模型已成为通用机器人操作的有前景范式,通过大规模预训练实现优异性能。该领域快速发展,引入空间先验和多样化架构创新,但不同训练方法和实现细节使性能提升来源难以厘清。本文提出SimVLA,一个简化基准模型,旨在为VLA研究提供透明参考。通过严格解耦感知与控制、采用标准视觉语言骨干和轻量动作头,并标准化关键训练动态,我们证明最小化设计也能达到顶尖性能。尽管仅有0.5B参数,SimVLA在无机器人预训练的仿真基准上超越多百亿参数模型。同时在真实机器人上表现与pi0.5相当。结果表明SimVLA是稳健、可复现的基线,有助于未来架构创新的性能归因。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic manipulation, leveraging large-scale pre-training to achieve strong performance. The field has rapidly evolved with additional spatial priors and diverse architectural innovations. However, these advancements are often accompanied by varying training recipes and implementation details, which can make it challenging to disentangle the precise source of empirical gains. In this work, we introduce SimVLA, a streamlined baseline designed to establish a transparent reference point for VLA research. By strictly decoupling perception from control, using a standard vision-language backbone and a lightweight action head, and standardizing critical training dynamics, we demonstrate that a minimal design can achieve state-of-the-art performance. Despite having only 0.5B parameters, SimVLA outperforms multi-billion-parameter models on standard simulation benchmarks without robot pretraining. SimVLA also reaches on-par real-robot performance compared to pi0.5. Our results establish SimVLA as a robust, reproducible baseline that enables clear attribution of empirical gains to future architectural innovations. Website: https://frontierrobo.github.io/SimVLA

机器人操作VLA模型轻量设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。