arXiv:2609.04552cs.ROcs.AI2026-09

让机器人部署后自主学习新任务,还能记住旧技能。

Continual Field-Adaptive Models (CFAMs) for Post-Deployment Physical AI

  • 用冻结慢学组件+快速学习胶囊场实现无梯度在线更新。
  • 仅用40%数据达全量训练模型性能,测试成功率提升13.9个百分点。
  • 适合需持续进化且算力受限的野外机器人系统。

无人交互自主——机器在危险环境中替代人类并使用人类工具完成任务——仍是关键任务中的缺失能力。这些场景中训练数据稀缺、仅具备机载计算资源,但部署系统必须应对新情况而不丢失已有能力。我们提出持续场适应模型(CFAM),在实验室高效学习,并在部署后通过自主、无梯度、设备端更新持续学习。CFAM采用互补学习架构:冻结的慢学习组件包含三个皮层——感知(将多模态输入映射为3D地面几何)、推理(将任务分解为技能并评估结果)、动作(执行几何化技能);胶囊场以无梯度方式存储单次学习的胜任力胶囊。技能安装在实验室为少样本,在场为持续进行;开放世界新奇性不在研究范围。我们在五种实体平台(机械臂、四足机器人、人形机器人、无人机、越野车)上评估了CFAM。基线模型(pi0、CogACT、SpatialVLA)使用相同自研多实体数据集进行物理平台对比。CFAM仅用40%数据即达到全量预训练策略的性能水平,或减少2.5倍轨迹数量。测试时,对已验证近分布外案例的自主捕获使动作成功率提升13.9个百分点。序列仿真中,其反向迁移仅-0.5百分点,而LoRA为-11.4。因此,CFAM提供了有界形式的部署后物理智能:少样本技能学习、基于验证近分布外经验的自主场域增长,以及先验能力保留。

原文摘要 · Abstract (English)

Unattended interactive autonomy - machines that step into danger in place of humans and complete tasks with human tools - remains a missing capability in mission-critical operations. These domains offer scarce training data and only onboard compute, yet deployed systems must face novelty without erasing prior competence. We introduce Continual Field-Adaptive Models (CFAMs), which learn efficiently in the lab and continue learning after deployment through autonomous, gradient-free, on-device updates. CFAM uses a complementary learning architecture with a frozen slow-learning component and a fast-learning Capsule Field. The slow component contains three cortices: Sensor, which maps multimodal input into 3D-grounded geometry; Reasoning, which decomposes tasks into skills and evaluates outcomes; and Action, which executes geometric skills. The Capsule Field stores field learning one-shot and gradient-free as Competence Capsules. Skill installation is few-shot in the lab and continual in the field; open-world novelty is outside scope. We evaluate CFAM across five embodiments: manipulator, quadruped, humanoid, quadrotor, and off-road vehicle. Baselines (pi0, CogACT, SpatialVLA) use the same in-house multi-embodiment dataset for physical-platform comparisons. CFAM reaches the operating point of a standard policy trained on the full prior-training dataset using 40% of the data, or 2.5x fewer trajectories. At test time, autonomous capture of verified near-OOD cases improves action success by 13.9 percentage points. In sequential simulation, backward transfer is -0.5 percentage points versus -11.4 for LoRA. CFAM therefore provides a bounded form of post-deployment physical intelligence: few-shot skill learning, autonomous field growth from verified near-OOD experience, and retention of prior competence.

持续学习机器人在线更新少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。