arXiv:2505.08361cs.AI2025-05ICLR被引 4

用可组合的因果组件建模未知环境,提升强化学习泛化能力

Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement Learning

  • 将环境分解为可组合的因果组件,通过语言引导实现结构化建模
  • 在仿真与真实机器人任务中,对未见任务的泛化性能超越现有方法
  • 适合需要快速适应新环境的机器人控制场景

强化学习中的泛化仍是重大挑战,尤其在遭遇动态未知的新环境时。受人类组合推理启发——将已知组件重组以应对新情况——我们提出世界建模的可组合因果组件(WM3C)框架。该框架通过学习并利用可组合的因果动态,增强强化学习的泛化能力。不同于以往关注不变表示学习或元学习的方法,WM3C识别并利用可组合元素间的因果关系,促进对新任务的稳健适应。本方法引入语言作为组合性模态,将潜在空间分解为有意义的组件,并在弱假设下提供其唯一识别的理论保障。实际实现采用带互信息约束和自适应稀疏正则化的掩码自编码器,以捕捉高层语义信息并有效解耦转移动态。在数值模拟与真实世界机器人操作任务上的实验表明,WM3C显著优于现有方法,在识别潜在过程、提升策略学习及泛化到未见任务方面表现优异。

原文摘要 · Abstract (English)

Generalization in reinforcement learning (RL) remains a significant challenge, especially when agents encounter novel environments with unseen dynamics. Drawing inspiration from human compositional reasoning -- where known components are reconfigured to handle new situations -- we introduce World Modeling with Compositional Causal Components (WM3C). This novel framework enhances RL generalization by learning and leveraging compositional causal components. Unlike previous approaches focusing on invariant representation learning or meta-learning, WM3C identifies and utilizes causal dynamics among composable elements, facilitating robust adaptation to new tasks. Our approach integrates language as a compositional modality to decompose the latent space into meaningful components and provides theoretical guarantees for their unique identification under mild assumptions. Our practical implementation uses a masked autoencoder with mutual information constraints and adaptive sparsity regularization to capture high-level semantic information and effectively disentangle transition dynamics. Experiments on numerical simulations and real-world robotic manipulation tasks demonstrate that WM3C significantly outperforms existing methods in identifying latent processes, improving policy learning, and generalizing to unseen tasks.

强化学习因果建模泛化机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。