arXiv:2606.18785cs.LGcs.AI2026-06

提出首个可随时停止的多目标强化学习算法,高效找到最优解集。

Bayesian Anytime Pareto Set Identification for Multi-Objective Multi-Armed Bandits

论文配图:Bayesian Anytime Pareto Set Identification for Multi-Objective Multi-Armed Bandits
图 1 · 摘自论文原文
  • 基于贝叶斯方法的TTPFTS算法,可随时停止并输出当前最优解集。
  • 在分子发现任务中,成功探索超大规模分子库并找到高质量候选方案。
  • 新增置信度度量,可实时评估算法对最优解集的把握程度。

识别帕累托最优解是支持多目标决策的关键。本文提出首个针对帕累托解集识别问题的任意时间(anytime)多目标多臂赌博机算法——基于贝叶斯的顶二帕累托前沿汤普森采样(TTPFTS)。我们在合成环境中将TTPFTS与最先进的固定预算算法对比。随后,在一个具有挑战性的多目标分子发现场景中展示了其实际效用,能够高效探索超大规模按需合成分子库。此外,我们引入一种新型不确定性量化指标,用于估计算法对预测帕累托集的信心。实验表明该指标能有效代理真实性能,提供复杂场景下学习进展的稳健监控方法。最后,我们通过理论证明了算法的渐近正确性。

原文摘要 · Abstract (English)

Identifying Pareto optimal solutions is critical to support multi-objective decision-making. We introduce the first anytime Multi-Objective Multi-Armed Bandit algorithm for the Pareto Set Identification problem, taking a Bayesian approach: Top-Two Pareto Front Thompson Sampling (TTPFTS). We benchmark TTPFTS against state-of-the-art fixed-budget Pareto Set Identification algorithms on synthetic environments. Next, we demonstrate its practical utility in a challenging multi-objective molecular discovery setting by efficiently exploring an ultra-large synthesis-on-demand molecular library. Furthermore, we introduce a novel uncertainty quantification metric that estimates our algorithm's confidence in the predicted Pareto set. We demonstrate that this metric effectively proxies true performance, yielding a robust methodology for monitoring learning progress in complex settings. Finally, we complement these empirical findings with a theoretical proof of the algorithm's asymptotic correctness.

多目标优化贝叶斯方法强化学习分子发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。