arXiv:2608.15269cs.RO2026-08

让机器人更聪明地记忆:用压缩视觉+双曲空间存储经验,提升长期任务成功率。

Remember Smarter: Visual History Compressor and Hyperbolic Experience Space for Robotic Memory

论文配图:Remember Smarter: Visual History Compressor and Hyperbolic Experience Space for Robotic Memory
图 1 · 摘自论文原文
  • 用双向空间与因果时间Mamba压缩多视角视觉历史,保留关键信息。
  • 在双曲空间中结构化存储成功经验,实现高效检索与无阻滞调用。
  • 适配pi0模型后,真实机器人任务成功率从53.6%提升至70.6%。

长时序机器人策略需要紧凑访问近期观测并复用经验,而不扩展视觉-语言-动作(VLA)上下文。我们提出「记住更智能」(RS),一个即插即用模块,包含互补的视觉历史压缩与双曲经验记忆分支。其视觉分支使用双向空间Mamba和因果时间Mamba压缩多视图图像块历史,通过残差交叉注意力将结果记忆注入面向动作的隐藏状态,同时保持原始视觉语言模型(VLM)的视觉标记流不变。其经验分支将成功的最终层VLM状态存储在庞加莱变分自编码器(Poincare VAE)空间中,以层级方式组织,并异步将检索到的经验转化为测地线提示标记,不阻塞动作推理。在pi0上适配后,RS将LIBERO-Plus任务总成功率从53.6%提升至70.6%,并在评估记忆保留与经验利用的真实机器人实验中取得显著性能提升。

原文摘要 · Abstract (English)

Long-horizon robot policies require compact access to recent observations and reusable experience without expanding the vision-language-action (VLA) context. We introduce Remember Smarter (RS), a plug-and-play module with complementary visual-history and hyperbolic experience-memory branches. Its visual branch compresses multi-view patch histories using bidirectional spatial Mamba and causal temporal Mamba, then exposes the resulting memory to action-facing hidden states through residual cross-attention while leaving the VLM visual-token stream unchanged. Its experience branch stores successful final-layer VLM states in a Poincare VAE space, organizes them hierarchically, and asynchronously converts retrieved experience into geodesic prompt tokens without blocking action inference. When adapted to pi0, RS increases total success on LIBERO-Plus from 53.6% to 70.6% and achieves substantial performance gains in real-robot experiments designed to evaluate memory retention and experience utilization.

机器人记忆双曲空间视觉压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。