arXiv:2605.10993cs.RO2026-05被引 1

提出连续分层记忆框架,提升视觉语言动作模型长程任务表现

ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models

论文配图:ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用双曲自动编码器将隐状态映射到分层空间,构建语义记忆树
  • 在LIBERO-Long上实现12.8%成功率提升,显著增强组合泛化能力
  • 适合研究长程操作、持续学习与具身智能的学者参考

记忆容量是决定视觉语言动作(VLA)模型在长程操作任务中性能的关键因素。现有记忆增强架构主要依赖线性或平坦存储,缺乏对操作类别和层次结构的先验知识,导致经验检索效率低,限制了对未见长程任务组合的泛化能力。受人类经验层级组织启发,我们提出ECHO(经验整合与层级组织)框架,该框架在连续分层空间中运行。通过双曲自动编码器,将VLA隐藏状态映射至该空间,利用双曲度量和蕴含约束机制,将经验向量组织为支持自顶向下检索的语义记忆树。同时,背景整合机制通过几何插值与结构分裂持续优化记忆树,支持连续空间中的虚拟记忆合成。我们将ECHO集成至$π_0$基础模型,在LIBERO数据集及初步真实世界实验中验证其有效性,显著提升执行成功率:在LIBERO-Long上相较$π_0$基线绝对提升12.8%,并在跨套件未见长程任务中改善组合泛化能力。

原文摘要 · Abstract (English)

Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing memory-augmented architectures primarily rely on linear or flat storage, lacking structural priors for manipulation categories and hierarchical organization. This deficiency hinders efficient experience retrieval and limits generalization to unseen long-horizon task compositions. Inspired by the hierarchical organization of human experience, we propose ECHO (Experience Consolidation and Hierarchical Organization), a novel memory framework operating within a Continuous Hierarchical Space. By employing a hyperbolic autoencoder, ECHO maps VLA hidden states into this space. Leveraging hyperbolic metrics and entailment constraint mechanisms, experience vectors are organized into a semantic memory tree that supports efficient top-down retrieval. In parallel, a background consolidation mechanism continuously refines the memory tree through geometric interpolation and structural splitting, supporting virtual memory synthesis in the continuous space. We integrate ECHO into the $π_0$ foundation model. Evaluations on LIBERO and preliminary real-world experiments demonstrate the effectiveness of our approach, notably achieving a 12.8% absolute improvement in execution success rate over the $π_0$ baseline on LIBERO-Long, while improving compositional generalization on cross-suite unseen long-horizon tasks.

视觉语言动作记忆机制分层结构长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。