arXiv:2602.24231cs.LG2026-02中稿 · paper

平衡探索与利用,实现组合实验设计的最优权衡。

Adaptive Combinatorial Experimental Design: Pareto Optimality for Decision-Making and Inference

  • 基于帕累托最优构建自适应组合实验框架。
  • 算法在有限时间内同时优化遗憾值与估计误差。
  • 更丰富的反馈信息显著提升决策精度,适合多目标场景。

本文首次系统研究自适应组合实验设计,聚焦组合多臂老虎机(CMAB)中遗憾最小化与统计功效之间的权衡。减少遗憾需频繁利用高回报动作,而准确推断奖励差距则需充分探索次优动作。我们通过帕累托最优概念形式化该权衡,并建立CMAB中帕累托高效学习的等价条件。考虑全反馈和半反馈两种信息结构,分别提出MixCombKL和MixCombUCB算法。理论证明两算法均为帕累托最优,在有限时间内同时保证遗憾和臂间差距估计误差的上界。结果表明,更丰富的反馈显著收紧可达帕累托前沿,主要增益来自所提方法下估计精度的提升。整体成果为多目标决策中的自适应组合实验提供了原则性框架。

原文摘要 · Abstract (English)

In this paper, we provide the first investigation into adaptive combinatorial experimental design, focusing on the trade-off between regret minimization and statistical power in combinatorial multi-armed bandits (CMAB). While minimizing regret requires repeated exploitation of high-reward arms, accurate inference on reward gaps requires sufficient exploration of suboptimal actions. We formalize this trade-off through the concept of Pareto optimality and establish equivalent conditions for Pareto-efficient learning in CMAB. We consider two relevant cases under different information structures, i.e., full-bandit feedback and semi-bandit feedback, and propose two algorithms MixCombKL and MixCombUCB respectively for these two cases. We provide theoretical guarantees showing that both algorithms are Pareto optimal, achieving finite-time guarantees on both regret and estimation error of arm gaps. Our results further reveal that richer feedback significantly tightens the attainable Pareto frontier, with the primary gains arising from improved estimation accuracy under our proposed methods. Taken together, these findings establish a principled framework for adaptive combinatorial experimentation in multi-objective decision-making.

实验设计多臂老虎机帕累托优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。