arXiv:2502.03725cs.LG2025-02

用机器学习求解流体非平稳多臂赌博机问题,高效生成可提速2600万倍的控制策略。

Optimal Control of Fluid Restless Multi-armed Bandits: A Machine Learning Approach

  • 基于状态方程为仿射或二次的特性,构建高效训练集
  • 通过非线性变换增强特征,提升策略学习效果
  • 训练出的时间相关状态反馈策略适用于维护、防疫等场景

我们提出一种新颖的机器学习框架,用于求解状态方程为仿射或二次的流体非平稳多臂赌博机问题(FRMABPs)。通过揭示FRMABPs的基本性质,我们设计了一种高效数值算法,通过求解多个具有不同初值的实例生成全面的训练集。进一步地,通过对特征向量施加非线性变换,利用FRMABPs的结构特性增强训练集。随后,采用带有超平面分割的最优分类树(OCT-H)学习时间依赖的状态反馈策略。我们在设备维护、疫情控制和渔业管理问题上测试该方法,结果表明所提方法能生成高质量的状态反馈策略。此外,一旦策略训练完成,其执行速度相比直接数值算法最高可提升2600万倍。

原文摘要 · Abstract (English)

We present a novel machine learning framework for the optimal control of fluid restless multi-armed bandit problems (FRMABPs) with state equations that are either affine or quadratic in the state variables. By establishing fundamental properties of FRMABPs, we develop an efficient numerical algorithm that generates a comprehensive training set by solving multiple instances with diverse initial states. We further enhance this training set by applying a nonlinear transformation to the feature vectors, leveraging structural properties of FRMABPs. A time-dependent state feedback policy is then learned using Optimal Classification Trees with hyperplane splits (OCT-H). We test our approach on machine maintenance, epidemic control, and fisheries control problems, demonstrating that our method yields high-quality state feedback policies. Furthermore, once a policy is learned, it achieves a speed-up of up to 26 million times compared to the direct numerical algorithm.

强化学习最优控制机器学习多臂赌博机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。