arXiv:2505.03912cs.ROcs.CV2025-05综述被引 89

开源双系统视觉语言动作模型,助力机器人操作研究

OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation

  • 对比分析现有双系统架构设计结构
  • 系统评估核心设计元素性能表现
  • 提供低成本可复用的开源模型

双系统视觉语言动作(VLA)架构已成为具身智能研究热点,但缺乏足够的开源工作以支持后续性能分析与优化。为解决此问题,本文总结并比较了现有双系统架构的结构设计,并对现有双系统架构的核心设计元素进行了系统性实证评估。最终,提供一个低成本开源模型,供进一步探索使用。该项目将持续更新更多实验结论及性能提升的开源模型,供各界选用。项目主页:https://openhelix-robot.github.io/

原文摘要 · Abstract (English)

Dual-system VLA (Vision-Language-Action) architectures have become a hot topic in embodied intelligence research, but there is a lack of sufficient open-source work for further performance analysis and optimization. To address this problem, this paper will summarize and compare the structural designs of existing dual-system architectures, and conduct systematic empirical evaluations on the core design elements of existing dual-system architectures. Ultimately, it will provide a low-cost open-source model for further exploration. Of course, this project will continue to update with more experimental conclusions and open-source models with improved performance for everyone to choose from. Project page: https://openhelix-robot.github.io/.

双系统VLA机器人操作开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。