arXiv:2410.08850math.OCcs.LG2024-10ICML被引 2

用深度学习解决大规模群体最优停止问题,实现高效可扩展的决策。

Learning to Stop: Deep Learning for Mean Field Optimal Stopping

  • 基于平均场理论构建连续群体最优停止模型,简化多智能体复杂度。
  • 在300维空间中验证方法有效性,6个问题上均实现高精度与低计算开销。
  • 适合研究金融、机器人等大规模协同决策场景的学者和工程师。

最优停止是优化中的基础问题,广泛应用于风险管理、金融、机器人及机器学习领域。本文将标准框架拓展至多智能体设置,称为多智能体最优停止(MAOS),其中多个智能体在有限状态、离散时间环境中协作做出最优停止决策。当智能体数量极大时,求解MAOS变得计算不可行,因此我们研究了在智能体数趋于无穷时的平均场最优停止(MFOS)问题。我们证明了MFOS对MAOS具有良好的近似性,并基于平均场控制理论建立了动态规划原理(DPP)。随后提出两种深度学习方法:一种通过模拟完整轨迹学习最优停止策略;另一种利用DPP反向递推求解值函数并学习最优停止规则。两种方法均训练神经网络以逼近最优停止策略。我们在6个不同问题上进行数值实验,空间维度最高达300,验证了方法的有效性与可扩展性。据我们所知,这是首个在离散时间与有限空间下形式化并计算求解MFOS的工作,为可扩展的多智能体最优停止方法开辟了新方向。

原文摘要 · Abstract (English)

Optimal stopping is a fundamental problem in optimization with applications in risk management, finance, robotics, and machine learning. We extend the standard framework to a multi-agent setting, named multi-agent optimal stopping (MAOS), where agents cooperate to make optimal stopping decisions in a finite-space, discrete-time environment. Since solving MAOS becomes computationally prohibitive as the number of agents is very large, we study the mean-field optimal stopping (MFOS) problem, obtained as the number of agents tends to infinity. We establish that MFOS provides a good approximation to MAOS and prove a dynamic programming principle (DPP) based on mean-field control theory. We then propose two deep learning approaches: one that learns optimal stopping decisions by simulating full trajectories and another that leverages the DPP to compute the value function and to learn the optimal stopping rule using backward induction. Both methods train neural networks to approximate optimal stopping policies. We demonstrate the effectiveness and the scalability of our work through numerical experiments on 6 different problems in spatial dimension up to 300. To the best of our knowledge, this is the first work to formalize and computationally solve MFOS in discrete time and finite space, opening new directions for scalable MAOS methods.

最优停止深度学习平均场多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。