arXiv:2606.20356math.OCcs.AI2026-06

针对共性噪声不确定的平均场控制,提出鲁棒Q学习算法。

Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise

  • 用量化投影+Wasserstein对偶重构共性噪声空间
  • 证明同步与异步算法收敛,给出有限步迭代上界
  • 在系统风险与流行病模型中验证鲁棒性与收敛性

本文针对离散时间平均场控制问题,在共性噪声分布存在Wasserstein不确定性的情形下,提出一种鲁棒Q-学习算法。该算法结合量化-投影策略与共性噪声空间上的Wasserstein对偶重表述。我们建立了同步与异步学习方案的收敛性,并给出了有限时间迭代上界。数值实验在系统性风险与流行病模型上进行,比较了异步实现与理想化贝尔曼迭代的表现,揭示了共性噪声误设下的鲁棒性-性能权衡,并报告了异步Q-学习算法的观测收敛行为。

原文摘要 · Abstract (English)

In this article, we present a robust $Q$-learning algorithm for discrete-time mean-field control problems under Wasserstein uncertainty in the common noise law. The algorithm combines a quantization-and-projection scheme with a Wasserstein dual reformulation on the common-noise space. We establish its convergence together with finite-time iteration bounds for both synchronous and asynchronous learning schemes. Numerical experiments on systemic risk and epidemic models compare the asynchronous implementation with an idealized Bellman iteration, illustrate the robustness-performance tradeoff under common-noise misspecification, and report the observed convergence behavior of the asynchronous $Q$-learning algorithm.

强化学习平均场控制鲁棒优化Wasserstein

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。