通过隐式类比实现跨情境任务组合,提升离线强化学习泛化能力
Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement Learning

- 用隐式类比表示法捕捉任务本质变化,忽略无关上下文差异
- 在OGBench上显著优于不支持类比推理的现有方法
- 适合需要在新场景中灵活完成未见目标的机器人控制任务
在离线目标条件强化学习中,泛化到未见过的目标与新情境至关重要,但受限于数据稀缺。现有方法依赖时间连续轨迹拼接,难以跨情境组合行为。本文提出将类比转导形式化为:基于已知类比与上下文合成新策略。所提类比表示能捕捉最优执行中的本质变化,对上下文变异保持不变,且足以实现最优目标达成。针对类比-上下文组合泛化难题,本文提出新方法,使离线GCRL可在未见组合上进行类比转导。在OGBench操作环境上的实验证明,该方法显著优于不支持类比转导的基线。
原文摘要 · Abstract (English)
Compositional generalization is essential for reaching unseen goals under novel contextual variations in offline goal-conditioned reinforcement learning (GCRL), where a generalist goal-reaching agent must be learned from limited data. Most prior approaches pursue this via trajectory stitching over temporally contiguous segments, which limits composing behaviors across varying contexts. To overcome this limitation, we formalize analogy transduction as synthesizing new plans by composing task-endogenous analogies with given contexts and propose a novel analogy representation tailored for it. Grounded in our theory, this analogy representation captures what changes under optimal task execution, remains invariant to contextual variations, and is sufficient for optimal goal reaching. We further contend that generalization to unseen analogy-context pairs is a practical obstacle in analogy transduction, and introduce a new approach for offline GCRL that enables analogy transduction beyond seen pairs to unseen combinations. We empirically demonstrate the effectiveness of our approach on OGBench manipulation environments, substantially outperforming prior methods that do not perform analogy transduction. Project page: https://rllab-snu.github.io/projects/CTA/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。