arXiv:2605.04895cs.LGstat.ML2026-05被引 1

提出条件化评估框架,解决多场景贝叶斯优化中排名不稳定问题。

Regime-Conditioned Evaluation in Multi-Context Bayesian Optimization

论文配图:Regime-Conditioned Evaluation in Multi-Context Bayesian Optimization
图 1 · 摘自论文原文
  • 基于预算比与先验相关性构建可迁移的制度评分PRS
  • 98%论文忽略预算比影响,实际排名常随预算反转
  • 适合关注优化结果稳定性的工业界研究者使用

现有迁移贝叶斯优化(transfer-BO)比较通常估计隐藏制度变量下的平均处理效应,而实际应用需要特定先验质量、预算比例和指标下的条件效应。对2022–2025年NeurIPS、ICML、ICLR等会议共40篇transfer-BO论文的审计显示,98%未将B/|A|作为控制变量。在相同GDSC2基准上,仅改变预算即反转排名:当B=50时,贪心策略优于UCB 0.050 Hit@1;当B=100时,UCB反超0.035。本文提出可移植制度评分PRS=(B/|A|)(1-rho),其中rho为先验秩相关性,可在主比较前从试点上下文估计。在涵盖化学、药物反应生物学和超参数优化(HPO)的79种条件下,分层模型得β=0.50(p=1.1e−9),19%条件处于优势小于0.01 Hit@1的等效区。在五个已发表的排名反转案例中,PRS能从前置可观测值预测胜者。提出“无免费排行榜”论点:当条件平均处理效应(CATE)在不同制度间符号变化时,报告的平均处理效应(ATE)会成为基准混合的函数。在线估计rho并动态切换策略的RegimePlanner,在B=100时于全部16个HPO搜索空间胜出,并在GDSC2上超越匹配的{Greedy, UCB}单上下文最优解18%。预注册预测整体准确率达67.5%(27/40),在EMA先验族内超过90%。实用协议建议:任何声称获取策略优势时,均需报告B/|A|、rho、K及指标。

原文摘要 · Abstract (English)

Published transfer-BO comparisons often estimate an average treatment effect of acquisition choice over hidden regime variables, while practitioners need the conditional effect for their specific prior quality, budget ratio, and metric. An audit of 40 transfer-BO papers from NeurIPS, ICML, ICLR, AISTATS, UAI, TMLR, JMLR, and AutoML-Conf (2022-2025) finds that 98% never vary B/|A| as a controlled axis. On the same GDSC2 benchmark, changing only the budget reverses the ranking: at B=50, Greedy outperforms UCB by 0.050 Hit@1, while at B=100, UCB outperforms Greedy by 0.035. We capture this transition with the Portable Regime Score PRS=(B/|A|)(1-rho), where rho is the prior rank correlation and can be estimated from pilot contexts before the main comparison. Across 79 conditions spanning chemistry, drug-response biology, and HPO, a hierarchical model gives beta=0.50 (p=1.1e-9), and 19% of conditions fall in an equivalence zone where |advantage|<0.01 Hit@1. In five published reversal cases, PRS predicts the winner from pre-comparison observables. A No-Free-Leaderboard proposition explains why unconditional rankings are unstable: when CATE changes sign across regimes, the reported ATE becomes a function of benchmark mixture. RegimePlanner, which estimates rho online and switches acquisition accordingly, wins all 16 HPO-B search spaces at B=100 and exceeds the matched {Greedy,UCB} per-context oracle on GDSC2 by 18%. Pre-registered predictions achieve 27/40=67.5% overall accuracy and above 90% within EMA prior families. The practical protocol is simple: report B/|A|, rho, K, and metric alongside any claimed acquisition advantage.

贝叶斯优化多场景评估制度条件超参调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。