arXiv:2508.19958cs.RO2025-08中稿 · CoRL被引 52

首个专为长时序机器人操作设计的视觉语言动作模型

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

  • 通过分阶段感知掩码,让模型自动区分移动与交互阶段
  • 在仿真和真实任务中显著超越现有方法,成功率提升超20%
  • 适合研究长时序机器人控制或想快速集成到现有模型的开发者

视觉-语言-动作(VLA)模型已成为机器人策略学习的核心,利用大规模多模态数据实现鲁棒且可扩展的控制。然而,现有VLA框架主要针对短时序任务,难以有效处理长时序、多步骤的机器人操作,受限于技能串联和子任务依赖问题。本文提出Long-VLA,首个专为长时序机器人任务设计的端到端VLA模型。其创新性地采用相位感知输入掩码策略,将每个子任务自适应划分为运动与交互阶段,使模型聚焦于阶段相关的感知线索,增强子任务兼容性。该统一策略保持了VLA训练的可扩展性和数据效率,其架构无关模块可无缝集成至现有VLA模型。我们还提出了L-CALVIN基准,系统评估长时序操作能力。大量仿真与真实世界实验表明,Long-VLA显著优于现有最先进方法,建立了长时序机器人控制的新基准。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing VLA frameworks primarily address short-horizon tasks, and their effectiveness on long-horizon, multi-step robotic manipulation remains limited due to challenges in skill chaining and subtask dependencies. In this work, we introduce Long-VLA, the first end-to-end VLA model specifically designed for long-horizon robotic tasks. Our approach features a novel phase-aware input masking strategy that adaptively segments each subtask into moving and interaction phases, enabling the model to focus on phase-relevant sensory cues and enhancing subtask compatibility. This unified strategy preserves the scalability and data efficiency of VLA training, and our architecture-agnostic module can be seamlessly integrated into existing VLA models. We further propose the L-CALVIN benchmark to systematically evaluate long-horizon manipulation. Extensive experiments on both simulated and real-world tasks demonstrate that Long-VLA significantly outperforms prior state-of-the-art methods, establishing a new baseline for long-horizon robotic control.

机器人控制长时序任务视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。