arXiv:2603.24576cs.ROcs.AI2026-03被引 3

让机器人记住过去关键信息,精准预测未来动作

Chameleon: Control-Indexed Prospective Memory for Visuomotor Manipulation

  • 用控制索引记忆法分离保存不同事件,确保可区分性
  • 在真实机器人任务中实现80.8%决策成功率,超越现有模型
  • 适合需要长时记忆与延迟决策的机器人操控场景

机器人常在执行动作前很久就观察到决定性信息。例如在杯球游戏中,机器人先看到球藏在哪只杯下,再观察杯子移动,最后才需选择正确杯子。仅凭最终观察无法决策,正确动作依赖早期事件。我们称这种时间差为观察-动作延迟,使记忆成为面向策略的问题:策略必须区分相似历史、定位当前决策相关的过往事件,并将回忆转化为可执行状态。我们提出Chameleon,一个约6000万参数的视觉运动策略,具备控制索引的前瞻性记忆能力。它能书写具身事件记忆,保持可分离的历史,检索与控制相关的痕迹,并训练出前瞻性的工作状态。我们还构建了Camo-Dataset,一个真实机器人基准数据集,通过使决策场景视觉模糊,强制依赖早期观察进行推理。Chameleon在该数据集上决策成功率从22.5%提升至80.8%,端到端成功率从21.3%升至71.3%。在公开的长时记忆基准上,其在LIBERO-10上达87.1%±0.8%,MemoryBench上97.3%±4.5%,MIKASA-Robo上75.1%±1.4%,达到同规模模型最优,且超过多个更大规模的VLA基线。

原文摘要 · Abstract (English)

Robots often observe information that determines a future action long before that action is executed. In a shell game, for example, a robot first sees which cup hides the ball, watches the cups move, and only later needs to choose the correct cup. The final observation alone is not enough for a decision: the correct action depends on an earlier event. We refer to this temporal gap as observation-action delay. It makes memory a policy-facing problem: a policy must keep similar histories distinct, retrieve the past event relevant to the current decision, and convert that recall into an action-ready state. We call these requirements separability, addressability, and prospectiveness. We introduce Chameleon, a ~60M visuomotor policy for control-indexed prospective memory. Chameleon writes embodied event memory, preserves separable histories, retrieves control-relevant traces, and trains the resulting working state to be prospective. We also introduce Camo-Dataset, a real-robot benchmark that isolates observation-action delay by making the decision scene visually ambiguous, so the correct action must be inferred from earlier observations. Chameleon improves decision/end-to-end success on Camo-Dataset from 22.5%/21.3% to 80.8%/71.3%. On public long-horizon memory benchmarks, it achieves 87.1% +/- 0.8% on LIBERO-10, 97.3% +/- 4.5% on MemoryBench, and 75.1% +/- 1.4% on MIKASA-Robo, setting the state of the art for same-size models and exceeding multiple larger VLA baselines under the reported protocols. Probes and ablations show that Chameleon learns separable, addressable, and prospective memory, and that these properties drive its performance gains.

机器人学习长时记忆视觉运动前瞻性记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。