arXiv:2603.25138quant-phcs.AI2026-03被引 2

用强化学习优化量子系统控制,实现近最优能量提取。

Reinforcement learning for quantum processes with memory

  • 设计基于最大似然估计的量子强化学习算法
  • 累计后悔值随回合数呈√K量级增长
  • 适用于未知量子记忆系统的实时自适应调控

在强化学习中,智能体通过与环境的序列交互来最大化奖励,仅能获得部分、概率性的反馈。这导致探索与利用的根本权衡:智能体必须探索以学习隐藏动态,同时利用已有知识最大化目标。尽管经典情形已广泛研究,但应用于量子系统需处理由未知动力学演化的隐藏量子态。我们建立一个框架,其中环境维持一个由未知量子通道演化的隐藏量子记忆,智能体通过量子仪器进行序列干预。针对此设定,我们采用一种乐观最大似然估计算法,并将其扩展至连续动作空间,以建模一般的正算子值测度(POVM)。通过控制估计误差在量子通道和仪器中的传播,我们证明该策略的累计后悔值在K个回合内为$ ilde{ m O}( extstyle oot{2} rom{K})$。此外,通过归约到多臂量子老虎机问题,我们建立了信息论下界,表明该次线性缩放在多项式对数因子范围内严格最优。作为物理应用,我们考虑状态无关的能量提取。当从一系列由隐藏记忆关联的非独立同分布量子态中提取自由能时,对源的任何不确定性都会导致热力学耗散。在此设定中,数学上的后悔值精确量化了累计耗散。使用我们的自适应算法,智能体利用过往能量结果实时优化提取协议,实现次线性累计耗散,从而获得渐近零耗散率。

原文摘要 · Abstract (English)

In reinforcement learning, an agent interacts sequentially with an environment to maximize a reward, receiving only partial, probabilistic feedback. This creates a fundamental exploration-exploitation trade-off: the agent must explore to learn the hidden dynamics while exploiting this knowledge to maximize its target objective. While extensively studied classically, applying this framework to quantum systems requires dealing with hidden quantum states that evolve via unknown dynamics. We formalize this problem via a framework where the environment maintains a hidden quantum memory evolving via unknown quantum channels, and the agent intervenes sequentially using quantum instruments. For this setting, we adapt an optimistic maximum-likelihood estimation algorithm. We extend the analysis to continuous action spaces, allowing us to model general positive operator-valued measures (POVMs). By controlling the propagation of estimation errors through quantum channels and instruments, we prove that the cumulative regret of our strategy scales as $\widetilde{\mathcal{O}}(\sqrt{K})$ over $K$ episodes. Furthermore, via a reduction to the multi-armed quantum bandit problem, we establish information-theoretic lower bounds demonstrating that this sublinear scaling is strictly optimal up to polylogarithmic factors. As a physical application, we consider state-agnostic work extraction. When extracting free energy from a sequence of non-i.i.d. quantum states correlated by a hidden memory, any lack of knowledge about the source leads to thermodynamic dissipation. In our setting, the mathematical regret exactly quantifies this cumulative dissipation. Using our adaptive algorithm, the agent uses past energy outcomes to improve its extraction protocol on the fly, achieving sublinear cumulative dissipation, and, consequently, an asymptotically zero dissipation rate.

强化学习量子控制能量提取后悔分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。