arXiv:2606.00780cs.LGcs.AI2026-06

提出新框架,让智能体在离线环境下更稳定地适应新任务。

Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning

论文配图:Behavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement Learning
图 1 · 摘自论文原文
  • 用基于Transformer的世界模型提取与行为无关的任务特征
  • 在稀疏奖励场景下,性能超越现有方法且更稳定
  • 适合需要强泛化能力的离线元强化学习场景

离线元强化学习通过静态数据集结合离线效率与元学习适应性,实现对未见环境的泛化,但面临上下文与策略分布偏移的挑战。这些问题导致智能体难以适应在线环境,尤其在稀疏奖励设置下更为严重,常陷入固有模式困境,无法实现鲁棒泛化。本文提出一种新框架,将信息论任务表示学习与基于Transformer的随机世界模型相结合。该方法提取对行为策略不变的任务定义潜在变量,有效缓解上下文分布偏移。为应对策略偏移与模型利用问题,引入保守价值惩罚机制,防止策略利用模型误差,同时保持强适应能力。大量实验表明,该方法在分布外和稀疏奖励设置下均优于现有先进方法,展现出更优的稳定性与泛化性能。

原文摘要 · Abstract (English)

Offline meta-reinforcement learning leverages static datasets to enable agents to generalize to unseen environments by combining offline efficiency with meta-learning adaptability, yet it faces key challenges from context and policy distribution shifts. These issues hinder agents from adapting to online environments, and are further exacerbated under sparse-reward settings. As a result, agents often become trapped in an inherent pattern dilemma, failing to achieve robust generalization. In this work, we propose a novel framework that integrates information-theoretic task representation learning with a Transformer-based stochastic world model. Our approach extracts task-defining latent variables that are invariant to behavior policy, thereby effectively mitigating the context distribution shift. To further handle policy shift and model exploitation, we apply a conservative value penalty to imagination-based rollouts, preventing the policy from exploiting model inaccuracies while maintaining robust adaptation. Extensive evaluations demonstrate that our method outperforms state-of-the-art approaches, with superior stability and generalization under out-of-distribution and sparse-reward settings.

元强化学习离线学习世界模型任务表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。