提出新方法主动优化强化学习中的模型估计效率
$κ$-Explorer: A Unified Framework for Active Model Estimation in MDPs
- 设计可调节参数的统一目标函数,平衡探索难度与访问频率
- 算法在基准测试中显著优于现有探索策略,性能更优
- 适合需要高效建模的强化学习场景,尤其关注精准估计
在状态完全可观测的表格化马尔可夫决策过程(MDP)中,每条轨迹提供基于状态-动作对的转移分布样本。因此,模型估计精度取决于探索策略如何根据各转移分布的内在复杂性分配访问频次。基于最近的覆盖探索工作,我们引入一个参数化、可分解且凹化的目标函数族 $U_κ$,显式融合了内在估计复杂度与外在访问频次。曲率参数 $κ$ 统一处理平均情况与最坏情况下的估计误差目标。利用 $U_κ$ 梯度的闭式表达,我们提出 $κ$-Explorer,一种在状态-动作占据度上进行 Frank-Wolfe 风格优化的主动探索算法。$U_κ$ 的递减收益结构自然优先考虑未充分探索且方差高的转移,同时保持光滑性以支持高效优化。我们为 $κ$-Explorer 建立了紧致的遗憾边界,并进一步提出一种全在线、计算高效的替代算法用于实际应用。在基准 MDP 上的实验表明,$κ$-Explorer 相较于现有探索策略表现出更优性能。
原文摘要 · Abstract (English)
In tabular Markov decision processes (MDPs) with perfect state observability, each trajectory provides active samples from the transition distributions conditioned on state-action pairs. Consequently, accurate model estimation depends on how the exploration policy allocates visitation frequencies in accordance with the intrinsic complexity of each transition distribution. Building on recent work on coverage-based exploration, we introduce a parameterized family of decomposable and concave objective functions $U_κ$ that explicitly incorporate both intrinsic estimation complexity and extrinsic visitation frequency. Moreover, the curvature $κ$ provides a unified treatment of various global objectives, such as the average-case and worst-case estimation error objectives. Using the closed-form characterization of the gradient of $U_κ$, we propose $κ$-Explorer, an active exploration algorithm that performs Frank-Wolfe-style optimization over state-action occupancy measures. The diminishing-returns structure of $U_κ$ naturally prioritizes underexplored and high-variance transitions, while preserving smoothness properties that enable efficient optimization. We establish tight regret guarantees for $κ$-Explorer and further introduce a fully online and computationally efficient surrogate algorithm for practical use. Experiments on benchmark MDPs demonstrate that $κ$-Explorer provides superior performance compared to existing exploration strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。