用知识蒸馏分步训练量子强化学习,让小量子模型也能搞定视觉控制。
Staged Hybridisation for Visual Quantum Reinforcement Learning via Knowledge Distillation

- 先训经典教师模型,再将策略行为蒸馏到紧凑的量子或经典头中
- 浅层量子电路头在像素环境中实现有效视觉控制,性能接近教师模型
- 适合想尝试小规模量子策略但资源有限的研究者
视觉环境对量子强化学习(QRL)构成严峻挑战:高维观测、不稳定的强化学习优化以及受限的变分量子电路(VQC)难以协同训练。本文提出以知识蒸馏(KD)作为分阶段混合训练策略。不直接从像素端到端训练混合智能体,而是先训练一个经典视觉教师模型,冻结其编码器作为特征接口,并将教师策略行为蒸馏至紧凑的下游头部。这些头部可为经典或基于VQC,使小型量子兼容学生模型能在与紧凑经典控制相同的冻结表示下进行评估。我们在CartPole Pixels和Acrobot Pixels上验证了该流程。结果表明,分阶段知识蒸馏使浅层VQC头部在原本极难直接像素训练的场景中获得非平凡的视觉控制能力。角度编码的VQC头部保持接近教师的性能,而幅度编码头部则达到极致紧凑,代价是更强的脆弱性、更高的预算敏感性及更长的模拟时间。总体而言,分阶段知识蒸馏将视觉QRL重构为紧凑头部学习问题,为在标准端到端强化学习循环外训练小型量子兼容策略开辟了实用路径。
原文摘要 · Abstract (English)
Visual environments are a demanding setting for quantum reinforcement learning (QRL): high-dimensional observations, unstable RL optimisation, and constrained variational quantum circuits (VQCs) are difficult to train jointly. This paper studies knowledge distillation (KD) as a staged hybridisation strategy for visual QRL. Instead of training a hybrid visual agent end-to-end from pixels, we first train a classical visual teacher, freeze its encoder as a feature interface, and distil the teacher's policy behaviour into compact downstream heads. These heads can be classical or VQC-based, enabling small quantum-compatible students to be evaluated under the same frozen representation as compact classical controls. We evaluate the pipeline on CartPole Pixels and Acrobot Pixels. The results show that staged KD enables shallow VQC heads to acquire non-trivial visual-control behaviour in settings where direct pixel-based training would be substantially more difficult. Angle-encoded VQC heads retain near-teacher performance, while amplitude-encoded heads push compactness to an extreme regime, at the cost of greater fragility, stronger budget sensitivity, and higher simulation time. Overall, staged KD reframes visual QRL as a compact-head learning problem, opening a practical route for training small quantum-compatible policies outside the standard end-to-end RL loop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。