提出分层记忆框架,让机器人长时任务更智能可靠。
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control

- 分三层结构:执行器、哨兵(工作记忆)、规划器(长期策略)
- 在长程任务中成功率显著提升,支持自我修正知识
- 适合需要持续推理的机器人控制场景
当前视觉-语言-动作(VLA)模型在机器人操作中表现优异,但在需长期记忆与推理的非马尔可夫任务中常遇瓶颈,因其依赖即时观测。现有方案面临“频率-能力悖论”:强推理模型过慢,快速模型能力不足。为此,我们提出分层具身记忆框架HiMe,将具身智能解耦为高频执行器、哨兵(工作记忆)与规划器(长期策略)。同时引入基于跨模态语义模板的动态知识系统与主动管理机制,支持“增、改、删”操作,保持记忆可塑性。该设计有效调和实时执行与慢思考规划的矛盾,在长时任务中显著提升成功率。实验表明,该方法不仅优于扁平记忆基线,还具备根据人类偏好自纠正内部知识的新能力。
原文摘要 · Abstract (English)
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions face a ''frequency-competence paradox,'' where stronger reasoning models are too slow for real-time control, while faster models lack sufficient reasoning capabilities. To resolve this architectural misalignment, we propose HiMe, a Hierarchical Embodied Memory framework that decouples embodied intelligence into a high-frequency Executor for execution, a Sentry for working memory, and a Planner for long-term strategy. We also introduce a dynamic knowledge system based on cross-modal semantic schemas and active management mechanisms, allowing robots to maintain memory plasticity through ''Add, Update, and Delete'' operations. This hierarchical design effectively balances the conflict between real-time execution and slow thinking planning, significantly improving success rates in long-horizon tasks. Experiments demonstrate that this approach not only outperforms flat memory baselines but also exhibits the novel ability to self-correct its internal knowledge based on human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。