arXiv:2410.01739cs.AIcs.LG2024-10中稿 · ICML

让强化学习像人一样用概念和信念高效决策

Conceptual Belief-Informed Reinforcement Learning

  • 从环境信息中抽象出高阶概念,构建概率信念作为经验先验
  • 在多种算法上提升采样效率,连续与离散控制任务均表现更好
  • 适合追求高效学习的强化学习研究者与工业应用开发者

强化学习虽取得显著进展,但受限于低效性和不稳定性,依赖大量试错数据,难以有效利用过往经验指导决策。人类则能高效学习,源于对概念的抽象及结合不确定性与先验知识更新概率信念,这正是认知科学观察到的现象。受此启发,我们提出概念信念引导的强化学习(HI-RL),一种可直接嵌入现有强化学习框架的高效经验利用范式。该方法通过提取关键环境信息的高层次类别形成概念,并构建自适应的概念相关概率信念,作为经验先验以指导价值或策略更新。我们在DQN、PPO、SAC和TD3等主流值函数与策略基算法中集成HI-RL,结果表明其在离散与连续控制基准测试中均实现一致的样本效率与性能提升。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has achieved significant success but is hindered by inefficiency and instability, relying on large amounts of trial-and-error data and failing to efficiently use past experiences to guide decisions. However, humans achieve remarkably efficient learning from experience, attributed to abstracting concepts and updating associated probabilistic beliefs by integrating both uncertainty and prior knowledge, as observed by cognitive science. Inspired by this, we introduce Conceptual Belief-Informed Reinforcement Learning to emulate human intelligence (HI-RL), an efficient experience utilization paradigm that can be directly integrated into existing RL frameworks. HI-RL forms concepts by extracting high-level categories of critical environmental information and then constructs adaptive concept-associated probabilistic beliefs as experience priors to guide value or policy updates. We evaluate HI-RL by integrating it into various existing value- and policy-based algorithms (DQN, PPO, SAC, and TD3) and demonstrate consistent improvements in sample efficiency and performance across both discrete and continuous control benchmarks.

强化学习概念抽象信念更新样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。