arXiv:2509.14163quant-phcs.LG2025-09被引 1

用量子强化学习动态调节图像生成中的引导参数,提升质量并减少模型参数。

Quantum Reinforcement Learning-Guided Diffusion Model for Image Synthesis via Hybrid Quantum-Classical Generative Model Architectures

  • 设计混合量子-经典架构,用量子电路生成策略特征,动态调整去噪引导值。
  • 在CIFAR-10上实现更优感知质量(LPIPS、PSNR、SSIM),且参数量更低。
  • 适合关注量子机器学习与生成模型融合的科研人员或开发者。

扩散模型通常采用静态或启发式分类器无关引导(CFG)调度,难以适应不同去噪步骤和噪声条件。本文提出一种量子强化学习(QRL)控制器,动态调整每一步的CFG值。该控制器采用混合量子-经典演员-评论家架构:浅层变分量子电路(VQC)带环形纠缠生成策略特征,经紧凑多层感知机(MLP)映射为ΔCFG的高斯动作;经典评论家估计价值函数。策略通过近端策略优化(PPO)与广义优势估计(GAE)优化,奖励函数平衡分类置信度、感知改善与动作正则化。在CIFAR-10上的实验表明,本方法在提升感知质量(LPIPS、PSNR、SSIM)的同时,相比经典强化学习演员和固定调度减少了参数量。消融实验揭示了比特数与电路深度在精度与效率间的权衡,扩展评估验证了长扩散流程下的鲁棒生成能力。

原文摘要 · Abstract (English)

Diffusion models typically employ static or heuristic classifier-free guidance (CFG) schedules, which often fail to adapt across timesteps and noise conditions. In this work, we introduce a quantum reinforcement learning (QRL) controller that dynamically adjusts CFG at each denoising step. The controller adopts a hybrid quantum--classical actor--critic architecture: a shallow variational quantum circuit (VQC) with ring entanglement generates policy features, which are mapped by a compact multilayer perceptron (MLP) into Gaussian actions over $Δ$CFG, while a classical critic estimates value functions. The policy is optimized using Proximal Policy Optimization (PPO) with Generalized Advantage Estimation (GAE), guided by a reward that balances classification confidence, perceptual improvement, and action regularization. Experiments on CIFAR-10 demonstrate that our QRL policy improves perceptual quality (LPIPS, PSNR, SSIM) while reducing parameter count compared to classical RL actors and fixed schedules. Ablation studies on qubit number and circuit depth reveal trade-offs between accuracy and efficiency, and extended evaluations confirm robust generation under long diffusion schedules.

量子机器学习扩散模型生成模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。