arXiv:2501.04572eess.SYcs.LG2025-01

用后悔值分析梯度下降,连接在线学习与自适应控制

Regret Analysis: a control perspective

  • 以后悔值视角分析凸函数梯度下降算法
  • 证明在线控制中误差可收敛至紧集
  • 适合研究控制与学习交叉的学者

在线学习与模型参考自适应控制存在诸多交集,但在算法分析目标和评价指标上差异显著。自适应控制关注系统参数/状态始终有界,且实时误差趋近于零(或紧集);而在线学习则常用后悔值衡量性能,即算法累积损失与事后最优固定参数损失之差。两者假设也不同:自适应控制基于输入输出特性,针对固定误差模型;在线学习则针对特定损失函数类(如凸函数),并预先假设某些信号有界。本文通过凸函数梯度下降的后悔值分析,以及流式回归问题的控制视角分析,深入探讨二者差异,并讨论在线自适应控制这一新范式。

原文摘要 · Abstract (English)

Online learning and model reference adaptive control have many interesting intersections. One area where they differ however is in how the algorithms are analyzed and what objective or metric is used to discriminate "good" algorithms from "bad" algorithms. In adaptive control there are usually two objectives: 1) prove that all time varying parameters/states of the system are bounded, and 2) that the instantaneous error between the adaptively controlled system and a reference system converges to zero over time (or at least a compact set). For online learning the performance of algorithms is often characterized by the regret the algorithm incurs. Regret is defined as the cumulative loss (cost) over time from the online algorithm minus the cumulative loss (cost) of the single optimal fixed parameter choice in hindsight. Another significant difference between the two areas of research is with regard to the assumptions made in order to obtain said results. Adaptive control makes assumptions about the input-output properties of the control problem and derives solutions for a fixed error model or optimization task. In the online learning literature results are derived for classes of loss functions (i.e. convex) while a priori assuming certain signals are bounded. In this work we discuss these differences in detail through the regret based analysis of gradient descent for convex functions and the control based analysis of a streaming regression problem. We close with a discussion about the newly defined paradigm of online adaptive control.

在线学习自适应控制后悔分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。