用进化算法生成初始策略,提升工业连续控制中强化学习的稳定性与性能。
Evolutionary Warm-Starts for Reinforcement Learning in Industrial Continuous Control
- 用CMA-ES生成高质量演示轨迹,作为强化学习的预热初始化。
- 相比随机初始化,新方法使训练更稳定,性能显著提升。
- 适合工业界需可靠、可复现控制策略的场景。
强化学习在工业控制中应用仍有限,部分原因在于难以在真实条件下训练出可靠的智能体。本文研究进化策略如何支持此类场景,提出一种面向工业分拣任务的连续控制基准。采用CMA-ES算法生成高质量示范轨迹,用于暖启动强化学习智能体。实验表明,基于CMA-ES的初始化能显著提升训练稳定性和最终性能。此外,生成的示范轨迹提供了强参考性能基准,本身也具有研究价值。本研究验证了混合进化-强化学习方法的有效性,为未来复杂工业应用奠定了基础。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is still rarely applied in industrial control, partly due to the difficulty of training reliable agents for real-world conditions. This work investigates how evolution strategies can support RL in such settings by introducing a continuous-control adaptation of an industrial sorting benchmark. The CMA-ES algorithm is used to generate high-quality demonstrations that warm-start RL agents. Results show that CMA-ES-guided initialization significantly improves stability and performance. Furthermore, the demonstration trajectories generated with the CMA-ES provide a strong oracle reference performance level, which is of interest in its own right. The study delivers a focused proof of concept for hybrid evolutionary-RL approaches and a basis for future, more complex industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。