通过因果信息优先化提升强化学习采样效率
Causal Information Prioritization for Efficient Reinforcement Learning
- 基于因子化马尔可夫决策过程,挖掘状态与动作对奖励的因果关系
- 在39个任务中显著优于现有方法,实现更高效探索与学习
- 适合复杂环境下的目标导向强化学习研究者使用
当前强化学习方法常因盲目探索而存在样本效率低的问题,根源在于忽略了状态、动作与奖励之间的因果关系。尽管已有因果方法尝试解决此问题,但缺乏对奖励引导下状态与动作因果理解的扎实建模,限制了学习效率。为此,本文提出一种新方法——因果信息优先化(CIP),通过因子化马尔可夫决策过程推断状态与动作各维度对奖励的因果关系,实现因果信息的优先处理。具体而言,CIP利用状态与奖励间的因果关系进行反事实数据增强,优先关注高影响状态特征;同时引入因果感知的赋能学习目标,显著提升智能体在复杂环境中执行奖励导向动作的效率。为全面评估有效性,我们在5种不同连续控制环境中进行了39个任务的实验,涵盖基于像素和稀疏奖励的运动与操作技能学习。结果表明,CIP在多种场景下均持续优于现有强化学习方法。
原文摘要 · Abstract (English)
Current Reinforcement Learning (RL) methods often suffer from sample-inefficiency, resulting from blind exploration strategies that neglect causal relationships among states, actions, and rewards. Although recent causal approaches aim to address this problem, they lack grounded modeling of reward-guided causal understanding of states and actions for goal-orientation, thus impairing learning efficiency. To tackle this issue, we propose a novel method named Causal Information Prioritization (CIP) that improves sample efficiency by leveraging factored MDPs to infer causal relationships between different dimensions of states and actions with respect to rewards, enabling the prioritization of causal information. Specifically, CIP identifies and leverages causal relationships between states and rewards to execute counterfactual data augmentation to prioritize high-impact state features under the causal understanding of the environments. Moreover, CIP integrates a causality-aware empowerment learning objective, which significantly enhances the agent's execution of reward-guided actions for more efficient exploration in complex environments. To fully assess the effectiveness of CIP, we conduct extensive experiments across 39 tasks in 5 diverse continuous control environments, encompassing both locomotion and manipulation skills learning with pixel-based and sparse reward settings. Experimental results demonstrate that CIP consistently outperforms existing RL methods across a wide range of scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。