arXiv:2503.07877stat.MLcs.IT2025-03

考虑不同动作成本的最优成对探索,降低总学习成本。

Cost-Aware Optimal Pairwise Pure Exploration

  • 基于追踪与停止原则,设计可处理零成本动作的新算法
  • 理论证明算法渐近逼近最低成本下界
  • 适用于需控制成本的探索任务,如资源受限场景

纯探索是多臂赌博机(MAB)的基本问题之一。现有研究多集中于特定纯探索任务,缺乏对通用纯探索问题的系统性视角。本文提出一个通用框架,聚焦于识别目标动作对之间的成对关系。不同于以往仅优化停止时间(即样本复杂度)的研究,本文引入动作特异性成本,目标是最小化学习过程中的累计成本。在具有动作成本的成对纯探索通用框架下,推导出性能下界。进而提出新算法CAET(Cost-Aware Pairwise Exploration Task),其基于追踪与停止原则,创新性地处理可能为零的动作成本,构成极具挑战性的场景。理论分析表明,CAET的性能渐近逼近下界。进一步讨论了特殊情况,包括扩展至后悔最小化这一MAB的核心方向。实验结果在多种设置下验证了CAET的有效性与高效性。

原文摘要 · Abstract (English)

Pure exploration is one of the fundamental problems in multi-armed bandits (MAB). However, existing works mostly focus on specific pure exploration tasks, without a holistic view of the general pure exploration problem. This work fills this gap by introducing a versatile framework to study pure exploration, with a focus on identifying the pairwise relationships between targeted arm pairs. Moreover, unlike existing works that only optimize the stopping time (i.e., sample complexity), this work considers that arms are associated with potentially different costs and targets at optimizing the cumulative cost that occurred during learning. Under the general framework of pairwise pure exploration with arm-specific costs, a performance lower bound is derived. Then, a novel algorithm, termed CAET (Cost-Aware Pairwise Exploration Task), is proposed. CAET builds on the track-and-stop principle with a novel design to handle the arm-specific costs, which can potentially be zero and thus represent a very challenging case. Theoretical analyses prove that the performance of CAET approaches the lower bound asymptotically. Special cases are further discussed, including an extension to regret minimization, which is another major focus of MAB. The effectiveness and efficiency of CAET are also verified through experimental results under various settings.

多臂赌博机纯探索成本感知算法设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。