arXiv:2512.05946cs.AIcs.ET2025-12被引 1

量子增强版强化学习模型,显著优化人力调度资源分配

Variational Quantum Rainbow Deep Q-Network for Optimizing Resource Allocation Problem

  • 用量子电路替代传统网络,利用量子叠加与纠缠提升表示能力
  • 在4个基准上使任务完成时间减少26.8%,优于经典方法4.9%-13.4%
  • 适合研究量子机器学习与复杂调度问题的学者使用

资源分配因组合复杂性仍属NP难问题。尽管深度强化学习(DRL)方法如Rainbow DQN通过优先经验回放和分布头提升可扩展性,但经典函数逼近器限制了其表达能力。本文提出变分量子Rainbow DQN(VQR-DQN),将环形拓扑变分量子电路与Rainbow DQN结合,利用量子叠加与纠缠优势。将人力分配问题(HRAP)建模为马尔可夫决策过程(MDP),动作空间基于人员能力、事件日程与转移时间。在4个HRAP基准测试中,VQR-DQN相比随机基线实现26.8%的归一化完工时间降低,优于Double DQN与经典Rainbow DQN达4.9%-13.4%。性能提升与电路表达能力、纠缠度及策略质量的理论关联一致,验证了量子增强DRL在大规模资源分配中的潜力。代码开源:https://github.com/Analytics-Everywhere-Lab/qtrl/

原文摘要 · Abstract (English)

Resource allocation remains NP-hard due to combinatorial complexity. While deep reinforcement learning (DRL) methods, such as the Rainbow Deep Q-Network (DQN), improve scalability through prioritized replay and distributional heads, classical function approximators limit their representational power. We introduce Variational Quantum Rainbow DQN (VQR-DQN), which integrates ring-topology variational quantum circuits with Rainbow DQN to leverage quantum superposition and entanglement. We frame the human resource allocation problem (HRAP) as a Markov decision process (MDP) with combinatorial action spaces based on officer capabilities, event schedules, and transition times. On four HRAP benchmarks, VQR-DQN achieves 26.8% normalized makespan reduction versus random baselines and outperforms Double DQN and classical Rainbow DQN by 4.9-13.4%. These gains align with theoretical connections between circuit expressibility, entanglement, and policy quality, demonstrating the potential of quantum-enhanced DRL for large-scale resource allocation. Our implementation is available at: https://github.com/Analytics-Everywhere-Lab/qtrl/.

量子强化学习资源分配变分量子电路DQN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。