arXiv:2606.21929cs.LG2026-06

在有限评估预算下,基于偏好选择最优模型组合。

Selective Ensemble Based on Preference-Directed Multi-Objective Bandits

论文配图:Selective Ensemble Based on Preference-Directed Multi-Objective Bandits
图 1 · 摘自论文原文
  • 用多目标强化学习框架建模部分偏好的模型选择问题
  • 提出新算法,在多种指标下实现近似最优的决策性能
  • 适合需要平衡准确率、鲁棒性等多重目标的模型集成场景

现代机器学习系统中的选择性集成需在有限评估预算下挑选表现优异的模型候选,而下游任务通常仅对准确性、鲁棒性和推理能力等能力给出部分偏好。这一设定自然形成一个在部分指定线性偏好下的序列决策问题。我们将其形式化为偏好导向的多目标老虎机(PDMOB),其中可接受的权衡由多面体偏好锥表示。基于此,我们引入帕累托C-最优性,其涵盖标准帕累托最优和单权重标量化作为特例。随后提出偏好导向上置信界(PrefUCB)算法,通过保持方向性置信区间引导探索。我们分析了基于指标和差距加权的损失函数,建立了两种准则的实例相关对数边界,恢复了经典情形下关于时长T的最优对数依赖性。在大规模预训练模型选择与机构约束下的在线资产配置任务上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Selective ensemble for modern machine learning systems requires choosing promising model candidates under limited evaluation budgets, while downstream tasks often specify only partial preferences over capabilities such as accuracy, robustness, and reasoning. This setting naturally gives rise to a sequential decision problem under partially specified linear preferences. We formalize it as preference-directed multi-objective bandits (PDMOB), where admissible trade-offs are represented by a polyhedral preference cone. Based on this formulation, we introduce Pareto $C$-optimality, which recovers standard Pareto optimality and single-weight scalarization as special cases. We then propose the preference-directed upper confidence bound (PrefUCB) algorithm, which maintains directional confidence intervals to guide exploration. We analyze both indicator-based and gap-weighted regret, and establish instance-dependent logarithmic bounds for both criteria, recovering the optimal logarithmic dependence on the horizon $T$ in classical special cases. Experiments on large pre-trained model selective ensemble tasks and online asset allocation under institutional mandates validate the efficacy of our method.

模型集成多目标优化强化学习偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。