arXiv:2608.06729cs.ROcs.CV2026-08

让机器人记住世界和自身状态,解决视觉盲区与任务遗忘问题

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

论文配图:AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用双记忆架构持久记录环境与自身状态
  • 在真实长任务中成功率提升17.5%
  • 仅用单手摄像头就能实现顶尖表现

尽管视觉-语言-动作(VLA)模型推动了具身智能发展,但其固有的反应式范式在部分可观测和长时序任务中严重受限。当仅依赖手腕摄像头时,物体离开视野会导致感知遗忘,多步执行中也会出现任务进度遗忘。为此,我们提出AtlasVLA,一种从直接反应式操作转向主动推理的新框架,通过持续的全局世界-自我状态建模克服瓶颈。该框架采用双记忆结构:4D持续世界状态记忆将瞬时2D观测提升为全局更新的体素哈希空间状态,解决视觉盲区;自我的工作状态记忆则追踪历史自我状态与任务进展。通过将扩散Transformer(DiT)基于此联合世界-自我状态进行条件化,实现了鲁棒的空间推理。在LIBERO、RLBench及真实世界基准上的广泛评估表明,AtlasVLA仅使用手腕摄像头即达到当前最佳性能。显著优于多视角基线,在LIBERO-Long上成功率达9.4%绝对提升,在真实世界长任务中提升17.5%。

原文摘要 · Abstract (English)

While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.

具身智能视觉语言动作长期任务状态记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。