arXiv:2608.18330cs.LG2026-08

提出新方法判断动态集成何时能提升回归模型表现。

When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift

论文配图:When Does Dynamic Ensembling Pay Off? Diagnosing Regionwise Gains in Regression under Distribution Shift
图 1 · 摘自论文原文
  • 基于小样本目标域数据,评估动态加权是否比静态融合更优。
  • 在12组数据偏移实验中,预测准确率达Spearman相关系数0.98。
  • 适合关注模型鲁棒性与动态集成效果的科研与工程人员。

输入依赖的回归模型动态组合是否优于最优静态融合,取决于分布偏移情况,且部署前难以预知。本文提出 $\\[\widehat{D}_{\mathrm{CF5}}$,仅需少量目标域标注样本即可估计区域级凸组合相对于最优静态凸融合的交叉验证增益——即逐区域调整信任的实际收益。在12组固定的数据集偏移对(空间、时间、领域、特征聚类)中,$\\widehat{D}_{\mathrm{CF5}}$ 对实际区域级测试增益的预测达到数据集层面Spearman相关系数+0.98(95%置信区间[+0.83, +1.00];p=5×10⁻⁵),包含两个推翻预注册预期的案例。敏感性分析中16组结果相关系数为+0.83,而其他探测诊断方法最高仅+0.66。该差异凸显了区域信任重分配的核心作用:区域凸组合增益相关性为+0.98,而仿射校正后的平滑协变量堆叠相关性仅为+0.01。受控生成实验表明,动态增益源于偏移异质性与局部能力的交互,随偏移严重度上升,在128至256个探针标签间可实现。Probe-Validated Ensemble Selector 在保持静态凸融合底线的前提下,仅在置信下界超过底线时启用候选动态集成。在预注册前瞻性批次中,其在全部12次运行中匹配或超越静态基线;两次部署使测试风险分别降低11%和16%,而被拒候选若未被拦截,将导致损失超静态基准30倍以上。本文发布OpenRegShift,一个可复现的回归集成在分布偏移下的评估框架。

原文摘要 · Abstract (English)

Whether input-dependent ("dynamic") combination of a regression model pool beats the best static blend depends on the shift and is rarely known before deployment. Can a small labeled target-domain probe tell us when reallocating trust across regions of the input space will pay off? We answer this with $\widehat{D}_{\mathrm{CF5}}$, which estimates from the probe the cross-fitted gain of the regionwise convex combination over the best static convex blend: the realizable value of deciding, region by region, whom to trust. Across a frozen suite of 12 dataset-shift pairs (spatial, temporal, domain, feature-cluster), $\widehat{D}_{\mathrm{CF5}}$ predicts realized regionwise test gains with dataset-level Spearman $+0.98$ (95% CI $[+0.83, +1.00]$; $p=5\times10^{-5}$), including two cases overturning preregistered expectations. The relationship holds in a 16-pair sensitivity analysis (Spearman $+0.83$), whereas alternative probe diagnostics reach at most $+0.66$. This contrast isolates regional trust reallocation: correlation is $+0.98$ for regionwise-convex gain, but $+0.01$ for smooth covariate-dependent stacking after affine correction. A controlled generator shows dynamic gains arise from the interaction of shift heterogeneity and local competence, increase with shift severity, and become realizable between 128 and 256 probe labels in the tested grid. The Probe-Validated Ensemble Selector chooses among a static affine stacker and dynamic realizers, deploying a candidate only when a held-out lower confidence bound clears the static-convex floor. In a preregistered prospective batch, it matched or improved the floor in all 12 runs; two deployments reduced test risk by 11% and 16%, while the gate rejected a candidate whose un-gated deployment incurred $>30\times$ the static loss. We release OpenRegShift, a reproducible evaluation harness for regression ensembles under distribution shift.

集成学习分布偏移动态融合模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。