在未知扰动下,该研究精确求解了单变量自适应控制的最优性能极限。
Consistent Model Chasing Is Minimax Optimal: The Exact Value of Scalar Adversarial Adaptive Control under Large Parametric Uncertainty

- 采用一致模型追踪策略,在参数不确定性下实现鲁棒最优控制。
- 理论证明最优峰值响应为 $1 + Δ$,其中 $Δ$ 为参数不确定范围。
- 适用于对对抗性扰动有严格性能要求的控制系统设计场景。
本文精确求解了一个基本的自适应控制问题:针对标量系统 $x_{t+1} = ax_t + u_t + w_t$,初始状态 $x_0=0$,且扰动满足 $ olinebreakorall t, |w_t| olinebreak ext{≤} olinebreak1$,其中未知的极点 $a \in [-Δ, Δ]$ 的符号和大小均未知,$Δ$ 可任意大。尽管系统形式简单,但迄今为止,对于任意大小参数不确定性的自适应控制问题,尚无文献给出其在对抗性干扰下最小最坏峰值 $ olinebreak ext{‖}x olinebreak ext{‖}_\infty$ 的确切值;现有理论仅提供稳定性保证、增益界和后悔率,而非精确值。本文证明该值为 $γ^\star(Δ) = 1 + Δ$(对所有 $Δ > 0$)。其中常数项 $1$ 是扰动不可避免的代价,$Δ$ 则是因必须识别而产生的唯一不可避峰值代价。最优控制律为在集合一致性区间中点处实施确定性等价死区控制,属于鲁棒预言机 × 一致模型追踪架构的实例。该架构是强制性的,非仅为充分条件:令 $θ_t := -u_t/x_t$,可将任意因果控制器表示为预言机选择器的组合,而最优性迫使选择器在关键历史路径上选取中点。传统工具——无论是经典还是现代方法——均会量化失效:探测行为在收益前即受惩罚,承诺机制在小于扰动激励时致命,乐观策略退化为平局决策或至少付出两倍于最优代价的长期成本,而后悔率证书无法捕捉双向最坏峰值。最优控制律不含任何探索机制,学习过程完全被动。这些结果首次为一致模型追踪作为对抗性自适应控制设计原则提供了精确最优性证明。
原文摘要 · Abstract (English)
We solve exactly a fundamental problem of adaptive control against adversarial disturbances: regulate the scalar system $x_{t+1} = ax_t + u_t + w_t$, $x_0=0$, $\|w\|_\infty \le 1$, where the constant pole $a \in [-Δ, Δ]$ is unknown in sign and magnitude and $Δ$ is arbitrarily large. Elementary as the system looks, the least worst-case peak $\|x\|_\infty$ that a causal controller can guarantee against an adversarial pair $(a, w)$ (the value of this game) has, to our knowledge, never been determined for any adaptive control problem with parametric uncertainty of arbitrary size under this criterion; existing theory supplies stability certificates, gain bounds, and regret rates, not the value. That value is $γ^\star(Δ) = 1 + Δ$ for every $Δ>0$. The summand $1$ is the irreducible price of the disturbance, and $Δ$ the exact price of a single, unavoidable identification spike. The optimal policy is certainty-equivalent deadbeat control at the midpoint of the set-membership consistent interval, an instance of the robust oracle $\times$ consistent model chasing architecture. The architecture is forced, not merely sufficient: writing $θ_t := -u_t/x_t$ exhibits every causal controller as an oracle-selector composition, and optimality pins the selector to the midpoint at the critical histories. The standard tools, classical and modern, each fail quantifiably: probing is punished before it pays, commitment is fatal at sub-disturbance excitation once adaptation is necessary, optimism degenerates to tie-breaking or pays asymptotically at least twice the optimum, and regret certificates are blind to the worst-case peak in both directions. The optimal law contains no exploration mechanism, its learning purely passive. These results give the first exact optimality certificate for consistent model chasing as a design principle for adversarial adaptive control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。