arXiv:2602.22663cs.RO2026-02中稿 · ICRA被引 7

提出轻量级视觉语言动作模型,实现在消费级显卡上的端到端移动操作。

Rethinking the Practicality of Vision-language-action Model: A Comprehensive Benchmark and An Improved Baseline

  • 设计跨仿真与真实世界的综合评测基准CEBench,覆盖多样机器人形态。
  • 构建14.4k模拟轨迹与1.6k真实专家轨迹,支持高效训练与评估。
  • 提出无需昂贵预训练的轻量模型LLaVA-VLA,适配移动端部署。

视觉-语言-动作(VLA)模型已成为通用机器人智能体,但现有模型存在参数量过大、预训练成本高、泛化能力差等问题。为提升其实用性,本文提出一个涵盖仿真与真实世界多种机器人形态的综合性评测基准CEBench,包含14.4k条模拟轨迹和1.6k条真实世界专家标注轨迹。基于CEBench,研究了三方面关键实用性问题,据此提出轻量级模型LLaVA-VLA:采用紧凑的视觉语言模型骨干网络,结合多视角感知、本体感知标记化与动作分块机制,并通过后训练与微调两阶段策略避免高昂预训练。该模型扩展动作空间以统一导航与操作任务。实验表明,其在多种机器人形态上具备强泛化与多任务适应能力;真实世界移动操作实验首次实现端到端的完整流程。所有数据集、代码与模型权重将在论文接收后开源。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a generalist robotic agent. However, existing VLAs are hindered by excessive parameter scales, prohibitive pre-training requirements, and limited applicability to diverse embodiments. To improve the practicality of VLAs, we propose a comprehensive benchmark and an improved baseline. First, we propose CEBench, a new benchmark spanning diverse embodiments in both simulation and the real world with consideration of domain randomization. We collect 14.4k simulated trajectories and 1.6k real-world expert-curated trajectories to support training on CEBench. Second, using CEBench as our testbed, we study three critical aspects of VLAs' practicality and offer several key findings. Informed by these findings, we introduce LLaVA-VLA, a lightweight yet powerful VLA designed for practical deployment on consumer-grade GPUs. Architecturally, it integrates a compact VLM backbone with multi-view perception, proprioceptive tokenization, and action chunking. To eliminate reliance on costly pre-training, LLaVA-VLA adopts a two-stage training paradigm including post-training and fine-tuning. Furthermore, LLaVA-VLA extends the action space to unify navigation and manipulation. Experiments across embodiments demonstrate the capabilities of generalization and versatility of LLaVA-VLA , while real-world mobile manipulation experiments establish it as the first end-to-end VLA model for mobile manipulation. We will open-source all datasets, codes, and checkpoints upon acceptance to foster reproducibility and future research.

机器人视觉语言动作轻量化模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。