arXiv:2502.02516cs.LGcs.AI2025-02ICML被引 5

针对多奖励多策略评估,提出自适应探索方法提升样本效率。

Adaptive Exploration for Multi-Reward Multi-Policy Evaluation

  • 基于实例的下界设计自适应探索策略,降低采样成本。
  • 在表格环境实验中显著减少评估所需样本数。
  • 适合需要高效评估多个策略与奖励的强化学习研究者。

我们研究在线多奖励多策略折扣设定下的策略评估问题,需同时评估不同策略在多个奖励函数下的表现。采用(ε,δ)-PAC视角,在有限或凸奖励集上实现高置信度的ε-准确估计,这一设置此前未被研究。借鉴多奖励最优策略识别的已有工作,将MR-NaS探索机制改进为联合最小化不同策略与奖励集下的样本复杂度。方法利用一个实例相关的下界,揭示样本复杂度随价值偏差程度的变化规律,指导高效探索策略设计。尽管计算该下界涉及困难的非凸优化,我们提出了适用于有限与凸奖励集的高效凸近似方法。在表格领域上的实验验证了该自适应探索方案的有效性。

原文摘要 · Abstract (English)

We study the policy evaluation problem in an online multi-reward multi-policy discounted setting, where multiple reward functions must be evaluated simultaneously for different policies. We adopt an $(ε,δ)$-PAC perspective to achieve $ε$-accurate estimates with high confidence across finite or convex sets of rewards, a setting that has not been investigated in the literature. Building on prior work on Multi-Reward Best Policy Identification, we adapt the MR-NaS exploration scheme to jointly minimize sample complexity for evaluating different policies across different reward sets. Our approach leverages an instance-specific lower bound revealing how the sample complexity scales with a measure of value deviation, guiding the design of an efficient exploration policy. Although computing this bound entails a hard non-convex optimization, we propose an efficient convex approximation that holds for both finite and convex reward sets. Experiments in tabular domains demonstrate the effectiveness of this adaptive exploration scheme.

强化学习策略评估多奖励自适应探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。