arXiv:2502.09298cs.LG2025-02

利用信念空间凸性提升部分可观测强化学习性能

Convex Is Back: Solving Belief MDPs With Convexity-Informed Deep Reinforcement Learning

  • 在信念空间中引入凸性约束,设计硬/软两种实现方式
  • 在Tiger和RockSample任务上显著提升智能体性能与鲁棒性
  • 特别适合对泛化能力要求高的新场景部署

我们提出一种新型深度强化学习方法,将部分可观测马尔可夫决策过程(POMDP)中价值函数在信念空间的凸性特性融入训练。引入硬约束和软约束两种凸性实现方式,并在经典的Tiger和FieldVisionRockSample两个POMDP环境中,与标准DRL进行对比。结果表明,加入凸性信息能显著提升智能体表现,并增强对超参数的鲁棒性,尤其在分布外测试域上优势明显。代码已开源:https://github.com/Dakout/Convex_DRL。

原文摘要 · Abstract (English)

We present a novel method for Deep Reinforcement Learning (DRL), incorporating the convex property of the value function over the belief space in Partially Observable Markov Decision Processes (POMDPs). We introduce hard- and soft-enforced convexity as two different approaches, and compare their performance against standard DRL on two well-known POMDP environments, namely the Tiger and FieldVisionRockSample problems. Our findings show that including the convexity feature can substantially increase performance of the agents, as well as increase robustness over the hyperparameter space, especially when testing on out-of-distribution domains. The source code for this work can be found at https://github.com/Dakout/Convex_DRL.

强化学习凸优化POMDP信念空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。