通过隐空间拓扑映射实现零样本模仿学习,让智能体在未见任务中成功适应。
Zero-shot Imitation Learning by Latent Topology Mapping

- 识别轨迹收敛/发散的隐空间枢纽状态,构建抽象转移拓扑
- 在3D迷宫中对未见任务实现55%零样本成功率,远超基线6%
- 适合需要快速泛化到新任务的复杂长时序控制场景
模仿学习在有专家演示时有效,但为每个复杂任务收集演示成本高昂。本文研究长时程、目标条件设置下,固定演示数据集包含有用行为但不覆盖所有任务的情况。现有方法虽能从演示中学习强策略,但在长时程任务中,微小误差沿动作轨迹累积,导致零样本迁移不可靠。为此提出零样本隐空间拓扑代理(ZALT),通过识别轨迹汇聚或发散的隐枢纽状态,学习枢纽间转移的策略与动态模型,并基于枢纽拓扑规划以完成新任务。该拓扑使演示行为显式可组合,同时将长任务压缩为更短的抽象转移序列,从而支持零样本适应。在复杂3D迷宫环境中,ZALT对未见任务达到55%的零样本成功率,显著优于最强基线的6%。
原文摘要 · Abstract (English)
Imitation learning is effective for training agents when expert demonstrations are available, but collecting demonstrations for every complex task in an environment is costly. We study the long-horizon, goal-conditioned setting where a fixed demonstration dataset contains useful behavior, but not complete examples for every task the agent must solve. Existing imitation learning methods can learn strong policies from demonstrations, but when solving long-horizon tasks, small errors accumulate over long primitive-action trajectories and make zero-shot adaptation to new tasks unreliable. We introduce Zero-shot Agents from Latent Topologies (ZALT), an imitation-learning method that solves unseen start-goal tasks beyond those demonstrated during training. ZALT identifies latent hub states where trajectories converge or diverge, learns policies and a dynamics model over hub-to-hub transitions, and plans over the hub topology to complete new tasks. This topology makes demonstrated behaviors explicitly composable while compressing long tasks into shorter sequences of abstract transitions -- combined, these enable ZALT to perform zero-shot adaptation. In a complex 3D maze environment, ZALT achieves 55% zero-shot success on unseen tasks, compared to 6% for the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。