arXiv:2607.02247math.STcs.LG2026-07

证明指数加权聚合在大温度下期望误差达到最优,解决长期悬而未决的问题。

Aggregation with Exponential Weights is Optimal in Expectation

  • 通过指数加权构造聚合估计器,无需强假设即可实现最优风险
  • 在温度满足特定条件时,期望超损失为 $T \log(M)/(n+1)$
  • 适用于模型选择中的平方损失,对预测值有界的情形特别有效

指数加权聚合(AEW)估计器在平方损失下的模型选择聚合基本设定中尚未完全理解。自引入以来,是否存在足够大的固定温度使AEW在期望意义下达到极小极大率最优,且在随机设计下仍成立,这一问题由Lecué和Mendelson(2013)明确提出,至今未解。本文证明:在不依赖伯恩斯坦型假设的前提下,只要温度 $T$ 满足 $(L^2/T)\ ext{exp}(B/T) \leq \mu/2$,AEW 的期望超损失即为 $T \log(M)/(n+1)$。其中 $M$ 为字典元素数量,$n$ 为独立同分布样本数,损失函数有界于 $B$,$L$-利普希茨连续且 $\mu$-强凸。对于平方损失,当预测值与标签均在 $[0, b]$ 范围内时,取 $T \geq 4b^2$ 即可。由于已知小温度下AEW在期望上次优,这表明其存在尖锐的相变行为,验证了Lecué与Mendelson的猜想。

原文摘要 · Abstract (English)

The aggregation with exponential weights (AEW) estimator is not fully understood in the basic setting of model selection aggregation with squared loss. In particular, whether it is minimax-rate optimal in expectation for large enough fixed temperatures and under random design has been an open problem since its introduction, which was explicitly posed by Lecué and Mendelson (2013). In this paper, we settle this problem by showing that \emph{without} requiring a Bernstein-type assumption, the AEW indeed achieves the excess risk $T \log (M) / (n+1)$ in expectation, whenever the temperature $T$ satisfies $(L^2/T)\exp(B/T)\leq μ/2$. Here, the number of dictionary elements is $M$, the estimator has observed $n$ i.i.d. samples from any distribution, and the loss is assumed to be bounded by $B$, $L$-Lipschitz continuous and $μ$-strongly convex. For squared loss, we show that $T\geq 4 b^2$ suffices when the predictions and labels are $[0,b]$-valued. Because AEW is known to be suboptimal in expectation for temperatures below some constant, this shows that AEW has a sharp phase transition when the temperature is large enough but constant, as conjectured by Lecué and Mendelson.

统计学习聚合估计指数加权

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。