arXiv:2511.12644cs.LG2025-11被引 1

重做20年前的强化学习算法,提升工业场景下的可复现性

NFQ2.0: The CartPole Benchmark Revisited

  • 用现代方法改进经典NFQ算法,聚焦批量学习与稳定性
  • 在真实工业组件系统上实现稳定且可复现的控制性能
  • 提供可操作的调参建议,助力工业强化学习落地

本文重新审视了20年前的经典神经拟合Q迭代(NFQ)算法在经典的CartPole基准上的表现。NFQ是早期将多层神经网络应用于现实控制问题的开创性工作,为现代深度强化学习奠定基础。尽管初始成功,但其依赖大量调参且难以在真实控制任务中复现。本文提出改进版本NFQ2.0,应用于基于标准工业组件构建的真实系统,重点提升学习过程的可复现性与鲁棒性。通过消融实验,揭示影响性能与稳定性的关键设计选择与超参数。最终表明,这些发现能帮助从业者更有效复现结果,并推动深度强化学习在工业场景中的实际应用。

原文摘要 · Abstract (English)

This article revisits the 20-year-old neural fitted Q-iteration (NFQ) algorithm on its classical CartPole benchmark. NFQ was a pioneering approach towards modern Deep Reinforcement Learning (Deep RL) in applying multi-layer neural networks to reinforcement learning for real-world control problems. We explore the algorithm's conceptual simplicity and its transition from online to batch learning, which contributed to its stability. Despite its initial success, NFQ required extensive tuning and was not easily reproducible on real-world control problems. We propose a modernized variant NFQ2.0 and apply it to the CartPole task, concentrating on a real-world system build from standard industrial components, to investigate and improve the learning process's repeatability and robustness. Through ablation studies, we highlight key design decisions and hyperparameters that enhance performance and stability of NFQ2.0 over the original variant. Finally, we demonstrate how our findings can assist practitioners in reproducing and improving results and applying deep reinforcement learning more effectively in industrial contexts.

强化学习可复现性工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。