arXiv:2605.31044cs.LG2026-05中稿 · Finding the Frame …

实测发现强化学习在工业供热系统中难落地,性能远低于仿真。

The Challenges of Using Reinforcement Learning for Controlling Industrial Energy Systems

  • 将控制问题建模为马尔可夫决策过程,分析可观测性、动作空间与奖励设计难题。
  • 真实系统中强化学习虽能保证稳定运行,但性能显著低于仿真结果。
  • 揭示了仿真到现实的差距,适合关注工业控制落地挑战的研究者。

强化学习在优化工业能源系统控制方面展现出巨大潜力,但现有研究多局限于仿真环境。本文以热力供热网络为案例,研究强化学习在真实工业系统部署中的挑战。将任务形式化为马尔可夫决策过程,系统分析其结构层面的难题:部分可观测性、动作空间设计、奖励函数设计以及仿真到现实的差距。这些挑战基于一项实际部署经验,结果显示强化学习在真实系统中实现了运行稳定性,但性能与仿真环境相比存在显著差距。

原文摘要 · Abstract (English)

Reinforcement learning has shown promising results for optimizing the control of industrial energy systems, yet most existing studies remain limited to the application in simulation environments. We investigate the challenges of deploying reinforcement learning in a real-world industrial energy system, considering a thermal heating network as a use case. We formulate the task as a Markov Decision Process and systematically analyze the associated challenges along the structure of the formal description, including partial observability, action space design, reward design, and the simulation-to-reality gap. The challenges are grounded in an existing real-world deployment, where reinforcement learning achieves operational stability but shows a significant performance gap compared to simulation.

强化学习工业控制仿真差距

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。