用代码代理管理记忆,让机器人更聪明地完成复杂操作。
Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation

- 结合代码代理与视觉语言动作模型,分层处理记忆与执行。
- 在RoboMemArena上任务成功率提升至60.1%,比基线高14.5个百分点。
- 只需示范基础动作,即可泛化到多种长时序任务,数据效率更高。
现代视觉-语言-动作(VLA)策略已掌握广泛的操作技能,但通常仅基于当前观测或固定长度历史生成动作。然而,真实世界操作常具非马尔可夫性,需机器人保留并推理长时程交互中的关键信息以决定下一步动作。为此,我们提出HyMeS混合学习框架,利用代码代理的推理与记忆管理能力,引导马尔可夫型VLA完成依赖记忆的操作。具体而言,HyMeS通过梯度仿射学习获得低级运动技能,而代码代理则通过迭代更新可执行启发式系统,从回放反馈中习得高级记忆管理策略。此外,通过多模态阶段完成验证闭环,利用本体感知信号和多帧视觉语言模型判断更新记忆。相较于端到端的记忆增强型VLA,HyMeS仅需为可复用的运动技能提供示范,而非每种依赖历史的任务配置,实现数据高效的组合泛化。在RoboMemArena上,HyMeS将平均累积成功率从52.5%提升至66.2%,平均任务成功率从41.3%提升至60.1%(对比pi0.5),累计成功率领先PrediMem 4.5个百分点,任务成功率领先14.5个百分点。
原文摘要 · Abstract (English)
Modern vision-language-action (VLA) policies have acquired broad manipulation skills, but typically generate each action chunk from the current observation or a short fixed-length history. However, real-world manipulation is often non-Markovian, requiring robots to retain and reason over task-relevant information from long-horizon interaction histories to determine the next action. To address this challenge, we propose HyMeS, a hybrid learning framework that leverages the reasoning and memory-management capabilities of coding agents to steer a Markovian VLA for memory-dependent manipulation. Specifically, HyMeS learns low-level motor skills through gradient-based imitation learning, while a coding agent acquires high-level memory-management strategies through heuristic learning by iteratively updating an executable heuristic system from rollout feedback. Furthermore, we close the loop between steering and execution through multimodal stage-completion verification, which updates memory using proprioceptive signals and multi-frame VLM judgments. Compared with end-to-end memory-augmented VLAs, HyMeS requires demonstrations only for reusable motor skills rather than for every history-dependent task configuration, enabling data-efficient compositional generalization. On RoboMemArena, HyMeS improves mean cumulative success from 52.5% to 66.2% and mean task success from 41.3% to 60.1% over pi0.5, while outperforming PrediMem by 4.5 points in cumulative success and 14.5 points in task success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。