arXiv:2505.17869cs.LG2025-05被引 2

在多目标强化学习中,找出最优的组合作,提升决策效率。

Best Group Identification in Multi-Objective Bandits

  • 基于淘汰机制设计算法,动态筛选最优组合作
  • 理论证明样本复杂度上界与下界,保证识别精度
  • 适用于多目标优化场景,如资源分配与推荐系统

我们提出了多目标多臂老虎机中的最佳组合作识别问题,其中智能体与具有向量奖励的臂组合进行交互。每个组的性能由一个效率向量表征,代表其在各维度上的最优可达成奖励。目标是在固定置信度设置下识别出最优组合作。研究了两种关键形式:组帕累托集识别(最优组的效率向量为帕累托最优)和线性最佳组合作识别(每维奖励有已知权重,最优组最大化效率向量加权和)。针对两种情形,我们提出基于淘汰的算法,建立了样本复杂度的上界,并推导出对任意正确算法均成立的下界。数值实验表明所提算法具有优异的实证性能。

原文摘要 · Abstract (English)

We introduce the Best Group Identification problem in a multi-objective multi-armed bandit setting, where an agent interacts with groups of arms with vector-valued rewards. The performance of a group is determined by an efficiency vector which represents the group's best attainable rewards across different dimensions. The objective is to identify the set of optimal groups in the fixed-confidence setting. We investigate two key formulations: group Pareto set identification, where efficiency vectors of optimal groups are Pareto optimal and linear best group identification, where each reward dimension has a known weight and the optimal group maximizes the weighted sum of its efficiency vector's entries. For both settings, we propose elimination-based algorithms, establish upper bounds on their sample complexity, and derive lower bounds that apply to any correct algorithm. Through numerical experiments, we demonstrate the strong empirical performance of the proposed algorithms.

多目标优化强化学习组合作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。