arXiv:2608.30880cs.RO2026-08

让机器人在操作中实时学习物理规律,越用越聪明。

Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation

论文配图:Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation
图 1 · 摘自论文原文
  • 用因果信号记录操作与变化,存入双时标记忆库。
  • 部署中成功率随经验积累持续提升,无需参数更新。
  • 适合需要自适应的现实场景机器人,如工业装配。

仅靠预训练难以实现可泛化的具身操作,因真实世界存在未见物理条件。我们主张机器人需在真实部署中通过自身物理交互实时学习,并据此指导后续动作。提出Zeva框架,首次实现机器人在保持策略模型冻结的前提下,基于自身交互经验进行上下文学习。Zeva采用因果交互提取器,将执行动作及其引发的状态变化编码为因果交互信号,并存储于双时标因果记忆中。后续动作时,从记忆中检索相关因果信号并注入冻结策略模型作为上下文。仿真与真实世界操作实验表明,Zeva在对比的前沿视觉语言模型(VLAs)和工作空间模型(WAMs)中表现最佳,更重要的是,其可在部署中实现自我进化,无需梯度更新。随着交互经验累积,成功率持续提升。且所获交互经验具备跨任务泛化能力。

原文摘要 · Abstract (English)

Generalizable embodied manipulation remains difficult to achieve through pretraining alone, due to unseen physical conditions in the real world. We argue that robots need to learn from their own physical interactions on the fly during real-world deployment and use this knowledge to inform subsequent actions. We present Zeva, the first framework that enables in-context learning from a robot's own physical interaction experience while keeping the policy model frozen. Zeva employs a Causal Interaction Extractor to encode an executed action and its induced state change into a causal interaction signal, which is stored in a dual-timescale causal memory. For subsequent actions, relevant causal interaction signals are retrieved from memory and injected into the frozen policy model as context. Experiments in simulation and real-world manipulation demonstrate that Zeva achieves the best performance among the compared frontier VLAs and WAMs and, more importantly, enables self-evolution during deployment without gradient updates. Its success rate continues to improve as the robot accumulates interaction experience. Furthermore, the acquired interaction experience can generalize across tasks.

具身智能因果学习自适应控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。