arXiv:2508.06659cs.LGcs.AI2025-08被引 2

通过可通信世界模型提升智能体零样本适应能力

In-Context Reinforcement Learning via Communicative World Models

  • 分离表征与控制,用信息代理生成简洁上下文消息
  • 在新任务中实现样本效率提升与零样本迁移
  • 适合需要快速适应新环境的强化学习场景

强化学习智能体通常难以在不更新参数的情况下泛化到新任务和新上下文,因其表征和策略过度拟合训练环境。为提升智能体的上下文内强化学习(ICRL)能力,本文将ICRL建模为双智能体涌现通信问题,提出CORAL(可通信自适应强化学习)框架。该框架通过功能分离隐式表征学习与控制决策,预训练一个信息代理(IA)作为世界模型,在多样化任务分布上学习世界建模并提炼为简洁消息。其通信协议由新颖的因果影响损失驱动,衡量消息对下一动作的影响。部署时,预训练的IA作为固定上下文提供者,辅助控制代理(CA)解码上下文并完成任务。实验表明,该方法显著提升样本效率,并在在线与离线环境中成功实现零样本适应,验证了可迁移通信表征的有效性。

原文摘要 · Abstract (English)

Reinforcement learning (RL) agents often struggle to generalize to new tasks and contexts without updating their parameters, mainly because their learned representations and policies are overfit to the specifics of their training environments. To boost agents' in-context RL (ICRL) ability, this work formulates ICRL as a two-agent emergent communication problem and introduces CORAL (Communicative Representation for Adaptive RL), a framework that learns a transferable communicative context by functionally separating latent representation learning from control. In CORAL, an Information Agent (IA) is pre-trained as a world model on a diverse distribution of tasks. Its objective is not direct return maximization, but world modeling and distilling its understanding into concise messages. The emergent communication protocol is shaped by a novel Causal Influence Loss, which measures the effect that the message has on the next action. During deployment, the previously trained IA serves as a fixed contextualizer for a new Control Agent (CA), which learns to solve tasks by interpreting the provided communicative context. Our experiments demonstrate that this approach enables the CA to achieve significant gains in sample efficiency and successfully perform zero-shot adaptation with the help of pre-trained IA in diverse online and offline environments, validating the efficacy of learning a transferable communicative representation.

强化学习零样本上下文学习通信机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。