量子深度强化学习让机器人更快学会复杂导航。
Quantum deep reinforcement learning for humanoid robot navigation task
- 用量子电路构建混合模型,直接处理高维状态空间。
- 量子SAC比经典SAC少92%步数,回报高8%(246.40 vs 228.36)。
- 适合研究量子增强机器人控制或高效强化学习的学者。
传统强化学习在高维、随机环境中因参数量大和非确定性而表现不佳。本文提出量子深度强化学习(QDRL),首次将该方法应用于拟人机器人,针对MuJoCo的Humanoid-v4和Walker2d-v4等具有庞大观测与动作空间的环境。通过参数化量子电路,构建混合量子-经典架构,直接在高维状态空间中实现导航,跳过传统映射与规划步骤。对比经典软演员-评论家(SAC)与量子版本,结果显示:量子SAC仅需92%的训练步数,平均回报达246.40,较经典SAC的228.36提升8%,证明量子计算在强化学习中具备加速潜力。
原文摘要 · Abstract (English)
Classical reinforcement learning (RL) methods often struggle in complex, high-dimensional environments because of their extensive parameter requirements and challenges posed by stochastic, non-deterministic settings. This study introduces quantum deep reinforcement learning (QDRL) to train humanoid agents efficiently. While previous quantum RL models focused on smaller environments, such as wheeled robots and robotic arms, our work pioneers the application of QDRL to humanoid robotics, specifically in environments with substantial observation and action spaces, such as MuJoCo's Humanoid-v4 and Walker2d-v4. Using parameterized quantum circuits, we explored a hybrid quantum-classical setup to directly navigate high-dimensional state spaces, bypassing traditional mapping and planning. By integrating quantum computing with deep RL, we aim to develop models that can efficiently learn complex navigation tasks in humanoid robots. We evaluated the performance of the Soft Actor-Critic (SAC) in classical RL against its quantum implementation. The results show that the quantum SAC achieves an 8% higher average return (246.40) than the classical SAC (228.36) after 92% fewer steps, highlighting the accelerated learning potential of quantum computing in RL tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。