arXiv:2502.21142cs.AIcs.LG2025-02被引 1

用大脑全局工作空间模型提升强化学习的想象力和鲁棒性

Multimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement Learning

  • 在高阶隐空间中进行思维模拟,替代直接处理像素
  • 训练所需环境步数减少,且缺失视觉或仿真属性仍保持稳定
  • 适合需要快速适应与多模态容错的智能体设计

人类利用对世界的丰富内在模型进行未来推理、想象反事实情景并灵活适应新情况。在强化学习中,世界模型旨在捕捉环境随智能体行为演变的规律,以支持规划与泛化。然而,传统世界模型直接作用于环境变量(如像素、物理属性),导致训练缓慢繁琐;相比之下,基于能捕捉相关多模态变量的高层隐空间可能更优。全局工作空间(GW)理论为大脑中的多模态整合与信息广播提供了认知框架,近期研究已开始引入高效的深度学习实现。本文评估了结合GW与世界模型的强化学习系统。将提出的GW-Dreamer与多种PPO及原始Dreamer版本对比,结果表明:在GW隐空间内执行思维模拟(即心智演练)可显著降低训练所需的环境步数。此外,该模型展现出一种涌现特性——即使缺失图像或仿真属性中的任一模态,仍具有强鲁棒性,而基线模型不具备此能力。结论表明,将全局工作空间与世界模型结合,极大提升了强化学习智能体的决策能力。

原文摘要 · Abstract (English)

Humans leverage rich internal models of the world to reason about the future, imagine counterfactuals, and adapt flexibly to new situations. In Reinforcement Learning (RL), world models aim to capture how the environment evolves in response to the agent's actions, facilitating planning and generalization. However, typical world models directly operate on the environment variables (e.g. pixels, physical attributes), which can make their training slow and cumbersome; instead, it may be advantageous to rely on high-level latent dimensions that capture relevant multimodal variables. Global Workspace (GW) Theory offers a cognitive framework for multimodal integration and information broadcasting in the brain, and recent studies have begun to introduce efficient deep learning implementations of GW. Here, we evaluate the capabilities of an RL system combining GW with a world model. We compare our GW-Dreamer with various versions of the standard PPO and the original Dreamer algorithms. We show that performing the dreaming process (i.e., mental simulation) inside the GW latent space allows for training with fewer environment steps. As an additional emergent property, the resulting model (but not its comparison baselines) displays strong robustness to the absence of one of its observation modalities (images or simulation attributes). We conclude that the combination of GW with World Models holds great potential for improving decision-making in RL agents.

强化学习世界模型多模态心智模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。