用自预测表征提升行为克隆在新状态-目标组合上的零样本泛化能力
Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning
- 通过自预测表征学习,让相似时间序列状态映射到相近隐空间
- 在多个需组合泛化的任务上达到竞争性性能,显著提升零样本效果
- 适合关注行为克隆泛化能力的研究者,尤其关注组合推理场景
虽然目标条件行为克隆(GCBC)方法在分布内训练任务上表现良好,但难以零样本泛化到需要对新状态-目标对进行条件控制的任务,即组合泛化。部分原因在于行为克隆学习的状态表征缺乏时间一致性:若时序相关的状态能被编码为相似的隐表示,则新状态-目标对的分布外差距将减小。本文通过证明,利用后续表示(Successor Representations, SR)促进长程时间一致性可助力泛化。进而提出一种简单有效的表示学习目标——BYOL-γ,理论上在有限马尔可夫决策过程(MDP)中近似SR,实证上在一系列需组合泛化的挑战性任务中表现出色。
原文摘要 · Abstract (English)
While goal-conditioned behavior cloning (GCBC) methods can perform well on in-distribution training tasks, they do not necessarily generalize zero-shot to tasks that require conditioning on novel state-goal pairs, i.e. combinatorial generalization. In part, this limitation can be attributed to a lack of temporal consistency in the state representation learned by BC; if temporally correlated states are properly encoded to similar latent representations, then the out-of-distribution gap for novel state-goal pairs would be reduced. We formalize this notion by demonstrating how encouraging long-range temporal consistency via successor representations (SR) can facilitate generalization. We then propose a simple yet effective representation learning objective, $\text{BYOL-}γ$ for GCBC, which theoretically approximates the successor representation in the finite MDP case through self-predictive representations, and achieves competitive empirical performance across a suite of challenging tasks requiring combinatorial generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。