LacaDM通过学习状态与策略间的潜在因果关系,提升多目标强化学习的适应性与泛化能力。
LacaDM: A Latent Causal Diffusion Model for Multiobjective Reinforcement Learning
- 在扩散模型中嵌入潜在时序因果结构,捕捉环境状态与策略的动态关联
- 在MOGymnasium上实现更高超体积、更低稀疏性与更优期望效用,超越现有基线
- 适用于复杂多目标场景,尤其适合需要跨任务迁移的智能体设计
多目标强化学习(MORL)因目标间固有冲突及动态环境适应难题而面临挑战。传统方法在大规模状态-动作空间中泛化能力不足。为此,我们提出潜变量因果扩散模型(LacaDM),用于增强离散与连续环境中MORL的适应性。不同于仅解决目标冲突的方法,LacaDM学习环境状态与策略之间的潜在时序因果关系,实现跨多样MORL场景的高效知识迁移。通过将此类因果结构嵌入基于扩散模型的框架,LacaDM在保持强泛化能力的同时,平衡了冲突目标。在MOGymnasium框架的多个任务上的实证评估表明,LacaDM在超体积、稀疏性和期望效用最大化方面持续优于最先进基线,展现出在复杂多目标任务中的有效性。
原文摘要 · Abstract (English)
Multiobjective reinforcement learning (MORL) poses significant challenges due to the inherent conflicts between objectives and the difficulty of adapting to dynamic environments. Traditional methods often struggle to generalize effectively, particularly in large and complex state-action spaces. To address these limitations, we introduce the Latent Causal Diffusion Model (LacaDM), a novel approach designed to enhance the adaptability of MORL in discrete and continuous environments. Unlike existing methods that primarily address conflicts between objectives, LacaDM learns latent temporal causal relationships between environmental states and policies, enabling efficient knowledge transfer across diverse MORL scenarios. By embedding these causal structures within a diffusion model-based framework, LacaDM achieves a balance between conflicting objectives while maintaining strong generalization capabilities in previously unseen environments. Empirical evaluations on various tasks from the MOGymnasium framework demonstrate that LacaDM consistently outperforms the state-of-art baselines in terms of hypervolume, sparsity, and expected utility maximization, showcasing its effectiveness in complex multiobjective tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。