让机器人更聪明地记事,用少量数据学会复杂操作。
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

- 引入时空记忆模块,让模型能记住过去观察和交互历史。
- 在多个任务和真实机器人上表现优异,且保持低数据需求。
- 适合需要长期记忆与泛化能力的开放世界机器人场景。
利用预训练视觉-语言模型(VLM)构建视觉-语言-动作(VLA)模型已成为3D机器人操作的有前景范式。然而,现有3D VLA方法仍存在数据依赖强、分布偏移下泛化能力弱、缺乏显式记忆等问题,限制了其在数据稀缺、开放世界和记忆依赖场景中的应用。此前的工作BridgeVLA通过保留预训练VLM的输入-输出对齐性,在3D动作学习中提升数据效率和泛化能力:将原始点云投影为多视角图像,预测中间热图后生成机器人动作。本文提出BridgeVLA++,在BridgeVLA基础上引入统一的时空记忆架构,建模持续的空间上下文与时间交互历史。该记忆增强框架可在保留原模型数据效率和泛化能力的基础上,实现对观测历史的推理。大量实验表明,该框架在空间操作任务中表现强劲,具备鲁棒泛化能力;在两个具有挑战性的记忆依赖操作基准上达到当前最优性能,且未牺牲原模型的数据效率与泛化性。此外,BridgeVLA++在双臂操作设置中表现良好,并在额外真实机器人平台上验证,证明其在任务、环境和平台间的可扩展性。这些结果确立BridgeVLA++为一个统一的3D视觉-语言-动作框架,同时支持数据高效学习、鲁棒泛化与有效的记忆感知机器人操作。
原文摘要 · Abstract (English)
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited generalization under distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world, and memory-dependent manipulation scenarios. Our previous work, BridgeVLA, improves data efficiency and generalization by preserving the input--output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions. In this work, we develop BridgeVLA++ by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities. Extensive experiments show that our framework achieves strong performance on spatial manipulation tasks while exhibiting robust generalization. BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ performs effectively in bimanual manipulation settings and is validated on an additional real-world robotic platform, demonstrating its scalability across tasks, environments, and robotic platforms. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization, and effective memory-aware robot manipulation. Project website: https://bridgevla-plus.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。