arXiv:2603.10126cs.ROcs.AI2026-03中稿 · ed被引 9

提出可长期记忆的自回归动作专家,让机器人更连贯地生成动作。

AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models

  • 用自回归序列生成动作,结合可更新的视觉语言前缀
  • 在模拟和真实机器人任务中成功率超主流反应式模型
  • 适合需要长时间上下文理解的智能机器人控制场景

我们提出一种独立的自回归(AR)动作专家,能够以连续因果序列生成动作,并基于可刷新的视觉-语言前缀进行条件约束。与现有视觉-语言-动作(VLA)模型及扩散策略不同,后者在每新观测时重置时间上下文并被动预测动作,我们的动作专家通过长期记忆维护自身历史,具备天然的上下文感知能力。该结构缓解了快速控制与缓慢推理之间的频率不匹配问题,支持运动学语法的独立高效预训练,并可模块化集成重型感知主干网络,自然保证跨帧时空一致的动作生成。为同步异步混合的视觉-语言-动作模态,我们采用重锚定机制,数学上处理训练与推理中感知信息的过时问题。在模拟和真实机器人操作任务上的实验表明,该方法能有效替代传统分块式动作头,适用于专用和通用策略。AR-VLA展现出更强的历史感知能力和显著更平滑的动作轨迹,同时保持或超越最先进反应式VLA的任务成功率。总体而言,本工作引入了一种可扩展、上下文感知的动作生成范式,为训练高效机器人策略提供了稳健的结构基础。代码与视频见 https://arvla.insait.ai

原文摘要 · Abstract (English)

We propose a standalone autoregressive (AR) Action Expert that generates actions as a continuous causal sequence while conditioning on refreshable vision-language prefixes. In contrast to existing Vision-Language-Action (VLA) models and diffusion policies that reset temporal context with each new observation and predict actions reactively, our Action Expert maintains its own history through a long-lived memory and is inherently context-aware. This structure addresses the frequency mismatch between fast control and slow reasoning, enabling efficient independent pretraining of kinematic syntax and modular integration with heavy perception backbones, naturally ensuring spatio-temporally consistent action generation across frames. To synchronize these asynchronous hybrid V-L-A modalities, we utilize a re-anchoring mechanism that mathematically accounts for perception staleness during both training and inference. Experiments on simulated and real-robot manipulation tasks demonstrate that the proposed method can effectively replace traditional chunk-based action heads for both specialist and generalist policies. AR-VLA exhibits superior history awareness and substantially smoother action trajectories while maintaining or exceeding the task success rates of state-of-the-art reactive VLAs. Overall, our work introduces a scalable, context-aware action generation schema that provides a robust structural foundation for training effective robotic policies. Code and Videos available at https://arvla.insait.ai

机器人控制自回归多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。