arXiv:2503.02735cs.LG2025-03被引 1

通过聚类目标策略优化重要性采样,提升多策略评估效率。

Clustered KL-barycenter design for policy evaluation

  • 将目标策略聚类后分别计算其KL中位点作为行为策略。
  • 理论证明样本复杂度上界优于单个中位点方法。
  • 适合需要高效评估多个策略的强化学习场景。

在随机老虎机模型中,本文研究如何设计高样本效率的行为策略,用于对多个目标策略的重要性采样评估。根据重要性采样理论,样本效率高度依赖于目标策略与重要性采样分布之间的KL散度。我们首先分析了以目标策略集合的KL中位点定义的单一行为策略;随后,通过将目标策略聚类为组内KL散度较小的子群,并为每组分配一个专属的KL中位点作为行为策略,进一步优化方法。该基于聚类的KL中位点策略评估(CKL-PE)算法提供了最优策略选择的新视角。我们推导了该方法的样本复杂度上界,并通过数值实验验证其有效性。

原文摘要 · Abstract (English)

In the context of stochastic bandit models, this article examines how to design sample-efficient behavior policies for the importance sampling evaluation of multiple target policies. From importance sampling theory, it is well established that sample efficiency is highly sensitive to the KL divergence between the target and importance sampling distributions. We first analyze a single behavior policy defined as the KL-barycenter of the target policies. Then, we refine this approach by clustering the target policies into groups with small KL divergences and assigning each cluster its own KL-barycenter as a behavior policy. This clustered KL-based policy evaluation (CKL-PE) algorithm provides a novel perspective on optimal policy selection. We prove upper bounds on the sample complexity of our method and demonstrate its effectiveness with numerical validation.

强化学习重要性采样策略评估聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。