用上下文引导的潜在世界模型提升离线元强化学习泛化能力
Contextual Latent World Models for Offline Meta Reinforcement Learning
- 将任务上下文嵌入潜在世界模型,实现任务相关的时间一致性建模
- 在多个基准上显著提升未见任务的泛化性能,尤其在复杂动力学场景中
- 适合研究离线强化学习、任务表示学习与自监督表征的学者
离线元强化学习旨在从固定数据集中学得可跨相关任务泛化的策略。基于上下文的方法通过转移历史推断任务表示,但无监督下学习有效任务表示仍具挑战。与此同时,潜在世界模型已通过时间一致性展现强大的自监督表征学习能力。本文提出上下文潜在世界模型,将潜在世界模型条件于推断出的任务表示,并与上下文编码器联合训练,强制任务条件下的时间一致性,从而获得捕捉任务依赖动态而非仅区分任务的表示。该方法学习到更具表达力的任务表示,在MuJoCo、Contextual-DeepMind Control和Meta-World三个基准上显著提升对未见任务的泛化能力。
原文摘要 · Abstract (English)
Offline meta-reinforcement learning seeks to learn policies that generalize across related tasks from fixed datasets. Context-based methods infer a task representation from transition histories, but learning effective task representations without supervision remains a challenge. In parallel, latent world models have demonstrated strong self-supervised representation learning through temporal consistency. We introduce contextual latent world models, which condition latent world models on inferred task representations and train them jointly with the context encoder. This enforces task-conditioned temporal consistency, yielding task representations that capture task-dependent dynamics rather than merely discriminating between tasks. Our method learns more expressive task representations and significantly improves generalization to unseen tasks across MuJoCo, Contextual-DeepMind Control, and Meta-World benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。