arXiv:2511.10187cs.LGcs.AI2025-11

用量子启发编码提升小样本强化学习性能

Improved Offline Reinforcement Learning via Quantum Metric Encoding

  • 用量子电路启发的编码器将状态映射到更紧凑的表示
  • 在三个各100样本的数据集上,性能提升超116%
  • 编码后状态空间几何更优,利于小样本训练

现实应用中强化学习常面临样本有限问题,现有离线RL方法表现不佳。本文提出量子度量编码器(QME),将原始状态嵌入更紧凑且有意义的表示空间,其结构受量子电路启发。对经典数据,QME为可经典模拟的可训练酉嵌入;对量子态数据,可直接在量子硬件上实现,无需测量或重编码。在三个各含100样本的数据集上,使用SAC与IQL算法验证,基于QME嵌入状态和解码奖励训练的智能体性能显著优于原始状态。平均而言,最大奖励性能相较提升116.2%(SAC)和117.6%(IQL)。进一步分析表明,QME嵌入后状态空间具有低Δ-双曲性,该几何特性有助于提升训练效率,为小样本下高效离线强化学习提供了新思路。

原文摘要 · Abstract (English)

Reinforcement learning (RL) with limited samples is common in real-world applications. However, offline RL performance under this constraint is often suboptimal. We consider an alternative approach to dealing with limited samples by introducing the Quantum Metric Encoder (QME). In this methodology, instead of applying the RL framework directly on the original states and rewards, we embed the states into a more compact and meaningful representation, where the structure of the encoding is inspired by quantum circuits. For classical data, QME is a classically simulable, trainable unitary embedding and thus serves as a quantum-inspired module, on a classical device. For quantum data in the form of quantum states, QME can be implemented directly on quantum hardware, allowing for training without measurement or re-encoding. We evaluated QME on three datasets, each limited to 100 samples. We use Soft-Actor-Critic (SAC) and Implicit-Q-Learning (IQL), two well-known RL algorithms, to demonstrate the effectiveness of our approach. From the experimental results, we find that training offline RL agents on QME-embedded states with decoded rewards yields significantly better performance than training on the original states and rewards. On average across the three datasets, for maximum reward performance, we achieve a 116.2% improvement for SAC and 117.6% for IQL. We further investigate the $Δ$-hyperbolicity of our framework, a geometric property of the state space known to be important for the RL training efficacy. The QME-embedded states exhibit low $Δ$-hyperbolicity, suggesting that the improvement after embedding arises from the modified geometry of the state space induced by QME. Thus, the low $Δ$-hyperbolicity and the corresponding effectiveness of QME could provide valuable information for developing efficient offline RL methods under limited-sample conditions.

强化学习小样本量子启发状态编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。