arXiv:2505.07607eess.SYcs.LG2025-05中稿 · DEXA 2025被引 3

用多目标强化学习让工业控制系统兼顾精准控制与省电。

Multi-Objective Reinforcement Learning for Energy-Efficient Industrial Control

  • 设计复合奖励函数,同时惩罚误差和耗电。
  • 当能耗权重低于0.25时,控制效果明显变差且非最优。
  • 适合关注工业节能与控制平衡的研究者或工程师。

工业自动化对节能控制策略的需求日益增长,需在性能、环境和成本之间取得平衡。本文针对单自由度的Quanser Aero 2测试平台,提出一种多目标强化学习(MORL)框架以实现节能控制。设计了复合奖励函数,同时惩罚跟踪误差和电能消耗。初步实验探究了能耗惩罚权重alpha在不同取值下对俯仰跟踪与节能之间权衡的影响。结果表明,当alpha在0.0至0.25之间时,系统性能出现显著变化;在仿真与真实系统中,较低的alpha值均导致非帕累托最优解。我们推测这可能源于Adam优化器自适应行为引入的偏差,可能偏好开关式控制策略。未来工作将通过高斯过程建模自动选择alpha,并推动该方法从仿真向实际部署过渡。

原文摘要 · Abstract (English)

Industrial automation increasingly demands energy-efficient control strategies to balance performance with environmental and cost constraints. In this work, we present a multi-objective reinforcement learning (MORL) framework for energy-efficient control of the Quanser Aero 2 testbed in its one-degree-of-freedom configuration. We design a composite reward function that simultaneously penalizes tracking error and electrical power consumption. Preliminary experiments explore the influence of varying the Energy penalty weight, alpha, on the trade-off between pitch tracking and energy savings. Our results reveal a marked performance shift for alpha values between 0.0 and 0.25, with non-Pareto optimal solutions emerging at lower alpha values, on both the simulation and the real system. We hypothesize that these effects may be attributed to artifacts introduced by the adaptive behavior of the Adam optimizer, which could bias the learning process and favor bang-bang control strategies. Future work will focus on automating alpha selection through Gaussian Process-based Pareto front modeling and transitioning the approach from simulation to real-world deployment.

强化学习工业控制节能优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。