arXiv:2410.23574math.OCcs.LG2024-10

提出新算法解决带记忆与有限预测的在线优化问题。

Online Convex Optimization with Memory and Limited Predictions

  • 设计可处理历史依赖与有限预测的新型算法
  • 动态后悔值随预测窗口长度指数下降
  • 适合需长期决策且有部分未来信息的场景

本文研究一种在线凸优化问题,其中每一步的代价函数依赖于过去决策的历史(即具有记忆性),且决策者仅能在有限窗口内获得未来代价值的有限预测。目标是设计算法以最小化相对于事后最优决策序列的动态后悔值。为此,我们提出一种新颖的预测算法,并建立了强理论保证:该算法的动态后悔值随预测窗口长度呈指数衰减。算法包含两个独立有价值的子模块:第一个解决带记忆和弱反馈的在线凸优化问题,实现√TV_T量级的动态后悔值,其中V_T衡量最优决策序列的变化;第二个为零阶方法,在一般凸优化中达到线性收敛速度,与一阶方法最优率相当。算法核心在于查询决策点时采用一种新颖的截断高斯平滑技术以获取预测。数值实验验证了理论结果。

原文摘要 · Abstract (English)

This paper addresses an online convex optimization problem where the cost function at each step depends on a history of past decisions (i.e., memory), and the decision maker has access to limited predictions of future cost values within a finite window. The goal is to design an algorithm that minimizes the dynamic regret against the optimal sequence of decisions in hindsight. To this end, we propose a novel predictive algorithm and establish strong theoretical guarantees for its performance. We show that the algorithm's dynamic regret decays exponentially with the length of the prediction window. Our algorithm comprises two general subroutines of independent interest. The first subroutine solves online convex optimization with memory and bandit feedback, achieving a $\sqrt{TV_T}$-dynamic regret, where $V_T$ measures the variation of the optimal decision sequence. The second is a zeroth-order method that attains a linear convergence rate for general convex optimization, matching the best achievable rate of first-order methods. The key to our algorithm is a novel truncated Gaussian smoothing technique when querying the decision points to obtain the predictions. We validate our theoretical results with numerical experiments.

在线优化动态后悔预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。