让机器人记住场景和任务经验,实现移动操作的精准执行。
EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation
- 引入情景与情节双记忆结构,模拟人类记忆机制
- 在仿真中实现0.52的成功率,比基线高0.20
- 适合需要导航与操作协同的移动机器人研究
视觉-语言-动作(VLA)模型虽已支持复杂任务理解,但多局限于短时程桌面操作,缺乏移动操作所需的记忆与推理能力。本文提出EchoVLA,一种面向移动操作的记忆感知型VLA模型。其融合受人脑启发的协同式声明性记忆:场景记忆维护空间语义地图,情节记忆存储带多模态上下文特征的任务经验。二者基于当前观测、任务历史与指令独立存取,通过粗细粒度注意力融合后指导基底机械臂扩散策略。为支持大规模训练,构建了MoMani自动化基准,通过多模态大语言模型规划生成专家级轨迹,并结合反馈优化,辅以真实机器人示范。仿真与真实世界实验表明,EchoVLA显著提升性能,在模拟环境中任务成功率达0.52(导航/操作),移动操作达0.31,分别优于强基线π_{0.5}的+0.20和+0.11。
原文摘要 · Abstract (English)
Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal instructions and perform complex tasks. However, existing VLAs are mostly confined to short-horizon, table-top manipulation, lacking the memory and reasoning capability required for mobile manipulation, where agents must coordinate navigation and manipulation under changing spatial contexts. In this work, we present EchoVLA, a memory-aware VLA model for mobile manipulation. EchoVLA incorporates a synergistic declarative memory inspired by the human brain, consisting of a scene memory that maintains a collection of spatial-semantic maps and an episodic memory that stores task-level experiences with multimodal contextual features. The two memories are individually stored, updated, and retrieved based on current observations, task history, and instructions, and their retrieved representations are fused via coarse- and fine-grained attention to guide base-arm diffusion policies. To support large-scale training, we further introduce MoMani, an automated benchmark that generates expert-level trajectories through multimodal large language model (MLLM)-guided planning and feedback-driven refinement, supplemented with real-robot demonstrations. Comprehensive simulated and real-world results demonstrate that EchoVLA substantially improves overall performance, e.g., it achieves the highest success rates of 0.52 on manipulation/navigation tasks and 0.31 on mobile manipulation tasks in simulation, exceeding the strong baseline $π_{0.5}$ by +0.20 and +0.11, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。