arXiv:2607.18866econ.EMcs.LG2026-07

用协方差视角重新定义后悔,找到最优策略方向。

Optimizing Regret

论文配图:Optimizing Regret
图 1 · 摘自论文原文
  • 将期望后悔转化为成本与决策的协方差,推导出其变分导数。
  • 线性策略下梯度为成本协方差矩阵,二阶导为零暗示边界最优解。
  • 适用于带约束优化,适合做在线学习与强化学习的理论研究者。

基于期望后悔等于成本与决策协方差的恒等式,本文发展了协方差后悔泛函的微分理论。推导出Gâteaux导数,表明普适的最陡下降方向是反向策略$-(c-\bar c)$,而上升方向对应动量。对于线性策略$\hatπ(c)=Ac+b$,其梯度为成本协方差矩阵$Σ_c$,零海森矩阵意味着边界最优解,如最小方差投资组合。研究扩展至约束优化,揭示后悔最小化与α最大化的符号梯度对偶性,给出了类Thompson采样的有限样本收敛界,并设计仅需观测输入的梯度下降算法。

原文摘要 · Abstract (English)

Building on the identity that expected regret equals the covariance between costs and decisions, this paper develops a derivative theory of the covariance regret functional. We derive the Gâteaux derivative, showing that the universal steepest-descent direction is the contrarian policy $-(c-\bar c)$, while ascent yields momentum. For linear policies $\hatπ(c)=Ac+b$, the gradient is the cost covariance matrix $Σ_c$, with a zero Hessian implying boundary-optimal solutions such as the minimum-variance portfolio. We extend to constrained optimization, sign-gradient duality between regret minimization and alpha maximization, finite-sample convergence bounds paralleling Thompson Sampling, and gradient-descent algorithms requiring only input observations.

后悔最小化在线学习凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。