arXiv:2511.12176quant-phcs.AI2025-11

用强化学习优化量子电池充电,实现在信息受限下的高效充能。

Reinforcement Learning for Charging Optimization of Inhomogeneous Dicke Quantum Batteries

  • 通过强化学习设计分段恒定充电策略,适应不同可观测条件。
  • 加入二阶关联信息后,充电性能达到全观测基准的94%-98%。
  • 策略具备前瞻性,牺牲短期波动以换取最终充能最优。

量子电池的充电优化是其实现的关键挑战,尤其在非均匀性和部分可观测性条件下。本文采用强化学习优化非均匀迪克模型电池的分段恒定充电策略,系统比较了四种可观测性场景:从全状态访问到实验可测的单个两能级系统(TLS)能量、一阶平均值及二阶相关性。仿真结果表明,全可观测性下可实现近最优的遍历能且波动极小;而在部分可观测性下,仅依赖单个TLS能量或能量加一阶平均值时,性能显著落后于全观测基线。然而,若补充二阶相关性信息,性能可恢复至全观测基线的94%-98%。所学充电计划具有非短视特性,为更优终端结果主动接受短暂平台期或下降。研究揭示了在现实信息约束下实现高效快充的可行路径。

原文摘要 · Abstract (English)

Charging optimization is a key challenge to the implementation of quantum batteries, particularly under inhomogeneity and partial observability. This paper employs reinforcement learning to optimize piecewise-constant charging policies for an inhomogeneous Dicke battery. We systematically compare policies across four observability regimes, from full-state access to experimentally accessible observables (energies of individual two-level systems (TLSs), first-order averages, and second-order correlations). Simulation results demonstrate that full observability yields near-optimal ergotropy with low variability, while under partial observability, access to only single-TLS energies or energies plus first-order averages lags behind the fully observed baseline. However, augmenting partial observations with second-order correlations recovers most of the gap, reaching 94%-98% of the full-state baseline. The learned schedules are nonmyopic, trading temporary plateaus or declines for superior terminal outcomes. These findings highlight a practical route to effective fast-charging protocols under realistic information constraints.

量子电池强化学习充电优化非均匀系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。