arXiv:2502.16772cs.LG2025-02ICML被引 1

提出新算法解决奖励不可见时的强化学习问题

Model-Based Exploration in Monitored Markov Decision Processes

  • 基于模型的区间估计,分两步分别处理可观测奖励与最优策略学习
  • 在40多个基准上收敛更快,已知监控机制时提升更显著
  • 首次给出有限样本性能保证,适用于部分奖励永远不可见场景

强化学习的一个基本假设是智能体始终能观测到奖励。但在许多现实场景中,这一假设不成立:例如人类观察者可能无法持续提供奖励,传感器受限或故障,或部署阶段奖励不可访问。为建模此类情况,最近提出了受监控马尔可夫决策过程(Mon-MDP)。然而现有算法存在多重局限:未能充分利用问题结构、无法利用已知监控器、对‘不可解’的Mon-MDP缺乏最坏情况保证且需特定初始化,仅提供渐近收敛证明。本文提出三项贡献:首先,设计一种新的基于模型的算法,克服上述缺陷。该算法采用双重模型化区间估计:一用于可靠捕获可观测奖励,另一用于学习极小极大最优策略。其次,通过实验证明其优势:在四十余个基准上收敛速度显著优于先前算法,尤其在已知监控过程时提升更为明显。第三,首次给出性能的有限样本界,证明即使某些奖励永远不可观测,算法仍可收敛至极小极大最优策略。

原文摘要 · Abstract (English)

A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or malfunctioning, or rewards may be inaccessible during deployment. Monitored Markov decision processes (Mon-MDPs) have recently been proposed to model such settings. However, existing Mon-MDP algorithms have several limitations: they do not fully exploit the problem structure, cannot leverage a known monitor, lack worst-case guarantees for 'unsolvable' Mon-MDPs without specific initialization, and offer only asymptotic convergence proofs. This paper makes three contributions. First, we introduce a model-based algorithm for Mon-MDPs that addresses these shortcomings. The algorithm employs two instances of model-based interval estimation: one to ensure that observable rewards are reliably captured, and another to learn the minimax-optimal policy. Second, we empirically demonstrate the advantages. We show faster convergence than prior algorithms in over four dozen benchmarks, and even more dramatic improvement when the monitoring process is known. Third, we present the first finite-sample bound on performance. We show convergence to a minimax-optimal policy even when some rewards are never observable.

强化学习模型算法有限样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。