用李普希茨正则化提升量子强化学习的鲁棒性与泛化能力
Robustness and Generalization in Quantum Reinforcement Learning via Lipschitz Regularization
- 通过李普希茨约束改进量子策略梯度算法,增强模型稳定性
- 实验显示新算法在训练中减少失败率,提升泛化性能
- 适合研究量子机器学习鲁棒性及强化学习优化的学者
量子机器学习相比经典方法具有更高精度和更低模型复杂度的潜力,有望推动多个领域的发展。在该领域中,量子强化学习常通过变分量子线路近似策略函数而实现。本文结合量子计算与控制理论,利用近期量子机器学习鲁棒性研究成果,提出基于李普希茨边界约束的正则化量子策略梯度方法(RegQPG)。实验表明,采用RegQPG训练能显著提升策略的鲁棒性与泛化能力。此外,我们还引入一种结合课程学习的算法变体,有效降低训练过程中的失败率。数值实验验证了该方法的实际优势。
原文摘要 · Abstract (English)
Quantum machine learning leverages quantum computing to enhance accuracy and reduce model complexity compared to classical approaches, promising significant advancements in various fields. Within this domain, quantum reinforcement learning has garnered attention, often realized using variational quantum circuits to approximate the policy function. This paper addresses the robustness and generalization of quantum reinforcement learning by combining principles from quantum computing and control theory. Leveraging recent results on robust quantum machine learning, we utilize Lipschitz bounds to propose a regularized version of a quantum policy gradient approach, named the RegQPG algorithm. We show that training with RegQPG improves the robustness and generalization of the resulting policies. Furthermore, we introduce an algorithmic variant that incorporates curriculum learning, which minimizes failures during training. Our findings are validated through numerical experiments, demonstrating the practical benefits of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。