arXiv:2606.29203cs.LGcs.IT2026-06

允许放弃推荐可让错误率从多项式降为指数级下降

Bayesian Best-Arm Identification with Abstention: A Polynomial-to-Exponential Phase Transition

论文配图:Bayesian Best-Arm Identification with Abstention: A Polynomial-to-Exponential Phase Transition
图 1 · 摘自论文原文
  • 引入放弃机制,通过控制放弃预算改变错误率衰减速度
  • 小放弃预算即可实现错误率指数下降,速率与放弃量平方成正比
  • 适用于追求高精度决策的贝叶斯场景,如医疗或金融选优

我们研究在采样预算固定下,允许学习者放弃推荐的贝叶斯最优臂识别问题。给定放弃预算 α,分析未检测到的错误概率——即未放弃却推荐次优臂的风险。核心发现:放弃机制引发相变——无放弃时错误概率随采样预算 T 多项式衰减;而引入任意小的正放弃预算后,衰减转为指数级。对于高斯先验和奖励,在 T→∞ 后 α↓0 的极限下,我们建立了信息论下界与算法上界精确匹配,最优错误指数形式为 exp(−α²T/(8κ_ν²))。其中 κ_ν 表示最优两臂差距在零处的先验密度,表明几乎持平的情形决定根本难度。我们提出自适应算法 PGWS,通过在统计模糊情形消耗放弃预算,实现了该最优指数。进一步证明此多项式到指数的提升仅是贝叶斯现象——在经典频率学派设定中,放弃仅影响低阶指数项。结果还扩展至高斯模型之外。

原文摘要 · Abstract (English)

We study the Bayesian fixed-budget best-arm identification problem in which a learner can abstain from making a terminal recommendation. Subject to an abstention budget $α$, we analyze the probability of undetected error--the risk of recommending a suboptimal arm without abstaining. Our central finding is that abstention induces a phase transition: without abstention, the error probability decays polynomially in the sampling budget $T$; in contrast, introducing any small positive abstention budget shifts this to an exponential decay. For Gaussian priors and rewards, in the regime $T\to\infty$ followed by $α\downarrow0$, we establish exact matching information-theoretic lower bounds and algorithmic upper bounds on the optimal error exponent, which takes the form $\exp(-\frac{α^{2}T}{8κ_ν^{2}})$. The hardness parameter $κ_ν$ represents the prior density of the top-two gap at zero, highlighting that nearly tied instances drive the fundamental error. We introduce an adaptive algorithm, PGWS, that successfully achieves this optimal exponent by expending its abstention budget on statistically ambiguous instances. We further demonstrate that this polynomial-to-exponential improvement is exclusively a Bayesian phenomenon--in the frequentist setting, abstention only affects lower-order exponent terms. We also extend our results beyond the Gaussian model.

贝叶斯优化主动学习相变最优臂识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。