arXiv:2510.23511cs.RO2025-10被引 15

开源视觉语言动作工具箱,一键复现主流智能体模型。

Dexbotic: Open-Source Vision-Language-Action Toolbox

  • 基于PyTorch的统一代码框架,支持多VLA策略并行实验。
  • 提供更强预训练模型,显著提升当前顶尖VLA性能。
  • 适合研究具身智能与多模态决策的开发者快速上手。

本文介绍Dexbotic,一个基于PyTorch的开源视觉-语言-动作(VLA)模型工具箱,旨在为具身智能领域的研究人员提供一站式服务。该工具箱支持多种主流VLA策略同时运行,用户仅需一次环境配置即可复现各类VLA方法。其以实验为中心的设计允许用户通过修改Exp脚本快速开展新实验。此外,我们提供了性能更优的预训练模型,显著提升了现有顶尖VLA策略的表现。Dexbotic将持续更新,集成更多最新的基础预训练模型与前沿VLA架构。

原文摘要 · Abstract (English)

In this paper, we present Dexbotic, an open-source Vision-Language-Action (VLA) model toolbox based on PyTorch. It aims to provide a one-stop VLA research service for professionals in the field of embodied intelligence. It offers a codebase that supports multiple mainstream VLA policies simultaneously, allowing users to reproduce various VLA methods with just a single environment setup. The toolbox is experiment-centric, where the users can quickly develop new VLA experiments by simply modifying the Exp script. Moreover, we provide much stronger pretrained models to achieve great performance improvements for state-of-the-art VLA policies. Dexbotic will continuously update to include more of the latest pre-trained foundation models and cutting-edge VLA models in the industry.

视觉语言动作具身智能开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。