arXiv:2603.12730cs.RO2026-03

用视觉锚点增强机器人操作中的空间与时间记忆能力。

AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation

  • 引入视觉锚点保留初始场景,结合轻量空间编码器捕捉几何关系。
  • 在模拟任务中成功率达80%,比基线提升13.6%。
  • 无需额外传感器,适合真实场景的低成本机器人控制应用。

当前视觉-语言-动作(VLA)系统在机器人操作中存在空间感知有限、缺乏历史记忆的问题。传统VLA仅基于单帧图像和语言指令生成动作,因2D图像无法提供精确空间信息,且无法保留过往上下文,导致物体被遮挡时易遗忘、操作中易迷失方向。为此,本文提出AnchorVLA4D,一种基于锚点的时空视觉-语言-动作模型。该模型通过添加锚定图像保持初始场景上下文,并引入轻量级空间编码器联合处理锚点与当前帧,显式揭示任务过程中的几何关系。模型基于Qwen2.5-VL主干网络并采用扩散动作头,无需深度或点云等额外传感模态,推理开销极小。结合冻结预训练空间编码器进一步提升性能,在Simpler WidowX基准上实现13.6%的提升,并在真实任务中达到平均80%的成功率。

原文摘要 · Abstract (English)

Since current Vision-Language-Action (VLA) systems suffer from limited spatial perception and the absence of memory throughout manipulation, we investigate visual anchors as a means to enhance spatial and temporal reasoning within VLA policies for robotic manipulation. Conventional VLAs generate actions by conditioning on a single current frame together with a language instruction. However, since the frame is encoded as a 2D image, it does not contain detailed spatial information, and the VLA similarly lacks any means to incorporate past context. As a result, it frequently forgets objects under occlusion and becomes spatially disoriented during the manipulation process. Thus, we propose AnchorVLA4D, a simple spatial-temporal VLA that augments the visual input with an anchor image to preserve the initial scene context throughout execution, and adds a lightweight spatial encoder that jointly processes the anchor and current frames to expose geometric relationships within an episode. Built on a Qwen2.5-VL backbone with a diffusion-based action head, AnchorVLA4D requires no additional sensing modalities (e.g., depth or point clouds) and introduces negligible inference overhead. Combining anchoring with a frozen pretrained spatial encoder yields further gains, realizing a 13.6% improvement on the Simpler WidowX benchmark and confirming the approach on real-world tasks, where it achieved an average success rate of 80%.

机器人操作视觉锚点时空建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。