一种新预测方法在多分类校准中同时达到最优后悔率,解决了长期存在的理论缺口。
Dirichlet Follow-the-Leader Closes the Gap in Simultaneous Multiclass U-Calibration
- 基于狄利克雷分布的贝叶斯自助法生成预测,无需调参
- 首次实现对所有有界损失和光滑损失的最优后悔率,上限为4√(KT)
- 适用于需要高精度校准的场景,如医疗诊断或金融风控
能否存在一个预测器,在任意有界正则损失下达到最优后悔率,并对任意光滑正则损失自适应?近期工作已接近答案,但仍有维度差距:其自共形扰动带来的最坏情况后悔率为约 $K^{5/4} oot{2}{T}$,对 $β$-光滑损失额外增加 $β oot{2}{K} \ \log K$。本文通过一个简洁的一行算法彻底填补这两项空白。观测到前 $t-1$ 步的类别计数 $c_{t-1}$ 后,从 $\ \operatorname{Dir}(c_{t-1})$ 中采样下一次预测,仅在已出现过的类别面上进行。这相当于对结果的全新贝叶斯自助。分析核心是一个精确恒等式:在 $\ \operatorname{Dir}(α)$ 下平均任一有界正则损失,等于其狄利克雷平均贝叶斯风险的离散导数。该恒等式使“成为被扰动领导者”项可望远镜化为非正的詹森差距。一个单计数似然比进一步将稳定性控制在该类别计数平方根的倒数内。最终得到的单一、无时域依赖的算法满足:对任意损失 $\ell$,$\sup_{\ell}\mathbb{E}\operatorname{Reg}_{\ell} \leq 4\sqrt{S_T T} \leq 4\sqrt{K T}$,且对任意 $β$-光滑正则损失,$\mathbb{E}\operatorname{Reg}_{\ell} \leq \frac{5}{2}β(1+\log T)$。其中 $S_T$ 为已观测类别数。已知下界表明这两个速率在其非平凡区域内均为最优。证明还覆盖了不可微损失及活动单纯形面的变化。
原文摘要 · Abstract (English)
Can one forecaster attain the optimal regret rate for every bounded proper loss and also adapt to every smooth proper loss? Recent work answered this up to a dimension gap. Its self-concordant perturbation gives roughly $K^{5/4}\sqrt{T}$ worst-case regret and incurs an additional $β\sqrt{K}\log K$ for $β$-smooth losses. We close both gaps with a one-line forecaster. After observing class counts $c_{t-1}$, draw the next prediction from $\operatorname{Dir}(c_{t-1})$, on the face of classes seen so far. This is a fresh Bayesian bootstrap of the outcomes. The analysis rests on an exact identity: averaging any bounded proper loss under $\operatorname{Dir}(α)$ equals a discrete derivative of its Dirichlet-averaged Bayes risk. The identity makes the be-the-perturbed-leader term telescope to a nonpositive Jensen gap. A one-count likelihood ratio then bounds stability by the inverse square root of that class's count. The resulting single, horizon-free algorithm satisfies $\sup_{\ell}\mathbb{E}\operatorname{Reg}_{\ell}\leq 4\sqrt{S_T T}\leq 4\sqrt{K T}$ and $\mathbb{E}\operatorname{Reg}_{\ell}\leq \frac{5}{2}β(1+\log T)$ for every $β$-smooth proper loss. Here $S_T$ is the number of observed classes. Known lower bounds show that both rates are optimal in their nontrivial regimes. The proof covers nondifferentiable losses and changes of the active simplex face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。