arXiv:2511.06252cs.LGcs.AI2025-11AAAI

让世界模型跨场景通用,提升强化学习的泛化能力。

MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios

论文配图:MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
图 1 · 摘自论文原文
  • 分解状态空间并引入元正则化,实现跨场景统一建模
  • 在多个仿真场景中显著优于现有最先进方法
  • 适合需要多任务泛化的强化学习研究者

基于模型的强化学习(MBRL)是提升强化学习算法泛化能力与样本效率的关键方法。然而,当前MBRL方法主要针对单一任务构建世界模型,很少关注跨不同场景的泛化能力。我们基于同一仿真引擎内动力学具有内在共性的观察,提出一种可跨场景泛化的统一世界模型——元正则化上下文世界模型(MrCoM)。该方法首先根据动态特性对潜在状态空间进行分解,提升世界模型预测精度;进一步采用元状态正则化提取场景相关性统一表示,并通过元值正则化使世界模型优化与策略学习在多样场景目标下保持对齐。我们理论分析了MrCoM在多场景设置下的泛化误差上界,并系统评估其在多种场景中的泛化性能,结果表明其显著优于现有最先进方法。

原文摘要 · Abstract (English)

Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different scenarios. Building on the insight that dynamics within the same simulation engine share inherent properties, we attempt to construct a unified world model capable of generalizing across different scenarios, named Meta-Regularized Contextual World-Model (MrCoM). This method first decomposes the latent state space into various components based on the dynamic characteristics, thereby enhancing the accuracy of world-model prediction. Further, MrCoM adopts meta-state regularization to extract unified representation of scenario-relevant information, and meta-value regularization to align world-model optimization with policy learning across diverse scenario objectives. We theoretically analyze the generalization error upper bound of MrCoM in multi-scenario settings. We systematically evaluate our algorithm's generalization ability across diverse scenarios, demonstrating significantly better performance than previous state-of-the-art methods.

强化学习世界模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。