用强化学习优化量子态测量,减少误差并提升效率
Bandits roaming Hilbert space
- 基于多臂老虎机框架在线优化量子观测选择
- 纯态下实现对数复杂度的后悔值,显著优于传统方法
- 适用于量子态层析、推荐系统及能量提取场景
本论文研究在在线学习量子态性质时的探索与利用权衡问题,采用多臂老虎机模型。在每一轮中,从一组可观测算符中选择一个以最大化其期望值,并利用历史信息优化策略以最小化累积后悔(即当前收益与最优可能收益之差)。推导了信息论下界并设计出匹配上界的最优策略,表明后悔值通常随轮次数的平方根增长。作为应用,将量子态层析重构为高效学习与最小化测量扰动的联合问题。针对纯态和连续动作空间,提出一种基于加权在线最小二乘估计器的样本最优算法,结合乐观原则控制设计矩阵特征值,实现多项式对数级别的后悔值。该框架还拓展至量子推荐系统与未知态下的热力学功提取。在后者中,结果表明其功耗耗散相比基于层析的方法具有指数级优势。
原文摘要 · Abstract (English)
This thesis studies the exploration and exploitation trade-off in online learning of properties of quantum states using multi-armed bandits. Given streaming access to an unknown quantum state, in each round we select an observable from a set of actions to maximize its expectation value. Using past information, we refine actions to minimize regret; the cumulative gap between current reward and the maximum possible. We derive information-theoretic lower bounds and optimal strategies with matching upper bounds, showing regret typically scales as the square root of rounds. As an application, we reframe quantum state tomography to both learn the state efficiently and minimize measurement disturbance. For pure states and continuous actions, we achieve polylogarithmic regret using a sample-optimal algorithm based on a weighted online least squares estimator. The algorithm relies on the optimistic principle and controls the eigenvalues of the design matrix. We also apply our framework to quantum recommender systems and thermodynamic work extraction from unknown states. In this last setting, our results demonstrate an exponential advantage in work dissipation over tomography-based protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。