提出自适应应力的更新机制,让模型在分布漂移下仍能稳定决策。
Stress-Aware Learning under KL Drift via Trust-Decayed Mirror Descent
- 引入应力感知的指数倾斜,同步优化信念与决策
- 实现$ ilde{O}( ext{sqrt}{T})$动态后悔,且每切换一次仅增$O(1)$代价
- 适合应对未知漂移、分布鲁棒性要求高的在线学习场景
我们研究分布漂移下的序列决策问题。提出熵正则化的信任衰减机制,将应力感知的指数倾斜同时注入信念更新与镜面下降决策中。在单纯形上,对偶等价表明信念倾斜与决策倾斜一致。通过脆弱性(KL球内最坏情况超额风险)、信念带宽(维持目标超额风险的半径)和决策空间脆弱性指数(在$O( ext{sqrt}{T})$后悔下容忍的漂移),形式化了鲁棒性。证明高概率敏感性边界,并在KL漂移路径长度$S_T = extstyle extsum_{t extge2} extsqrt{{ m KL}(D_t|D_{t-1})/2}$下建立$ ilde{O}( ext{sqrt}{T})$动态后悔保证。信任衰减实现每切换一次仅$O(1)$的后悔,而无应力更新存在$Ω(1)$尾部损失。一个无参数的博弈策略自适应未知漂移,持续过度倾斜则导致$Ω(λ^2 T)$静态惩罚。还获得校准应力边界,并扩展至二阶更新、弱反馈、异常值、应力变化、分布式优化及插件式KL漂移估计。该框架统一了动态后悔分析、分布鲁棒目标与KL正则化控制。
原文摘要 · Abstract (English)
We study sequential decision-making under distribution drift. We propose entropy-regularized trust-decay, which injects stress-aware exponential tilting into both belief updates and mirror-descent decisions. On the simplex, a Fenchel-dual equivalence shows that belief tilt and decision tilt coincide. We formalize robustness via fragility (worst-case excess risk in a KL ball), belief bandwidth (radius sustaining a target excess), and a decision-space Fragility Index (drift tolerated at $O(\sqrt{T})$ regret). We prove high-probability sensitivity bounds and establish dynamic-regret guarantees of $\tilde{O}(\sqrt{T})$ under KL-drift path length $S_T = \sum_{t\ge2}\sqrt{{\rm KL}(D_t|D_{t-1})/2}$. In particular, trust-decay achieves $O(1)$ per-switch regret, while stress-free updates incur $Ω(1)$ tails. A parameter-free hedge adapts the tilt to unknown drift, whereas persistent over-tilting yields an $Ω(λ^2 T)$ stationary penalty. We further obtain calibrated-stress bounds and extensions to second-order updates, bandit feedback, outliers, stress variation, distributed optimization, and plug-in KL-drift estimation. The framework unifies dynamic-regret analysis, distributionally robust objectives, and KL-regularized control within a single stress-adaptive update.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。