arXiv:2602.20403cs.LGmath.OC2026-02

提出在线鲁棒学习新框架,用最优传输度量保障决策在最坏分布下的稳定性。

Risk-Averse Wasserstein Distributionally Robust Online Learning

  • 将在线学习建模为决策者与对手的博弈,对手选择最坏分布
  • 算法收敛至与离线最优解一致的鲁棒纳什均衡
  • 适合追求高可靠性、抗分布偏移的实时决策场景

我们研究分布鲁棒在线学习,其中风险规避的学习者通过序列更新决策,以防范从过去观测值为中心的Wasserstein模糊集抽取的最坏分布。尽管该范式在离线设置下通过Wasserstein分布鲁棒优化(DRO)已得到充分理解,其在线扩展在收敛性方面仍面临重大挑战。本文将问题建模为决策者与选择最坏分布的对手之间的在线鞍点随机博弈,并提出一个通用框架,使其收敛到与相应离线Wasserstein DRO问题解一致的鲁棒纳什均衡。

原文摘要 · Abstract (English)

We study distributionally robust online learning, where a risk-averse learner updates decisions sequentially to guard against worst-case distributions drawn from a Wasserstein ambiguity set centered at past observations. While this paradigm is well understood in the offline setting through Wasserstein Distributionally Robust Optimization (DRO), its online extension poses significant challenges in convergence. In this paper, we formulate the problem as an online saddle-point stochastic game between a decision maker and an adversary selecting worst-case distributions, and propose a general framework that converges to a robust Nash equilibrium coinciding with the solution of the corresponding offline Wasserstein DRO problem.

在线学习鲁棒优化分布鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。