arXiv:2606.08028cs.LG2026-06中稿 · 2026 European Conf…被引 2

提出噪声自适应的在线凸优化高概率边界,提升稳定性与精度。

Noise-Adaptive High-Probability Regret Bounds for Online Convex Optimization

  • 用指数超鞅方法处理无界子高斯噪声,避免截断误差。
  • 带噪反馈下置信成本为log(1/δ)线性增长,优于全信息下的平方根。
  • 首次实现约束优化中后悔值与约束违反的同步高概率控制。

研究强凸损失下在线凸优化(OCO)的高概率后悔边界,解决噪声自适应、反馈结构与约束满足的开放问题。在全信息设置中,针对子高斯随机梯度,证明了噪声自适应的高概率后悔界:鞅偏离项随噪声水平σ而非梯度有界值G缩放,相比经典Azuma-Hoeffding界实现G/σ的乘法改进。分析引入指数超鞅论证,绕过Freedman不等式对有界差分的要求,可直接处理无界子高斯噪声而无需截断。在无信息反馈(bandit)场景中,证明最小最大下界:高概率后悔随log(1/δ)线性增长,而全信息下为√log(1/δ),形成反馈模型间置信成本的严格分离。对于满足Slater条件的随机约束OCO,同时给出累积后悔与长期约束违反的高概率保证,分别达到O(√T log(m/δ))和O(√T/(ζδ) + m√T log(m/δ))。合成实验验证所有理论预测。

原文摘要 · Abstract (English)

We study high-probability regret bounds for online convex optimization (OCO) with strongly convex losses and establish three results that resolve open questions at the intersection of noise adaptivity, feedback structure, and constraint satisfaction. For the full-information setting with sub-Gaussian stochastic gradients, we prove a noise-adaptive high-probability regret bound in which the martingale deviation term scales with the noise level $σ$ rather than the gradient bound $G$, yielding a multiplicative improvement of $G/σ$ over the classical Azuma-Hoeffding baseline. Our analysis introduces an exponential supermartingale argument that bypasses the bounded-difference requirement of Freedman's inequality, enabling direct treatment of unbounded sub-Gaussian noise without truncation artifacts. For bandit feedback, we prove a minimax lower bound: the high-probability regret scales linearly in $\log(1/δ)$, in contrast to the $\sqrt{\log(1/δ)}$ confidence cost under full information. This constitutes a formal separation in the confidence cost of strongly convex OCO across feedback models. Regarding constrained OCO with stochastic constraints satisfying a Slater condition, we provide simultaneous high-probability guarantees for both cumulative regret and long-run constraint violation, achieving $\mathcal{O}(\sqrt{T\log(m/δ)})$ regret and $\mathcal{O}(\sqrt{T}/(ζδ) + m\sqrt{T\log(m/δ)})$ violation. Synthetic experiments corroborate all theoretical predictions.

在线优化高概率界噪声自适应约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。