提出量子博弈中更合理的后悔度衡量方式,突破传统限制。
Coherent Swap Regret and Channel-Proof Learning
- 引入相干交换后悔度,衡量对局部量子操作的偏离
- 实现 $O( oot{d T log d})$ 的后悔率,优于经典方法
- 适用于量子相关均衡验证与探测带宽扩展场景
外部后悔仅衡量以固定策略替换自身行为的稳定性。在量子博弈中,这一标准忽略了自然的物理操作:玩家可对其实际接收或准备的状态施加局部完全正定保迹(CPTP)映射。本文提出相干交换后悔度,作为对所有此类局部CPTP偏差的后悔基准,并通过在CPTP Choi截面上采用熵镜面上升结合定点策略,实现了 $O( oot{d T log d})$ 的相干交换后悔。主要结果为三层偏差类别景观:替换通道恢复经典外部后悔率 $Θ( oot{T log d})$;幺模通道(包括酉操作及酉混合)最小最大后悔为零;确定性测量-制备通道在中等时长下即导致 $Ω( oot{d T log d})$ 后悔,且该速率足以覆盖所有CPTP偏差。因此,困难根源在于非幺模使用推荐寄存器,而非量子相干本身。应用上,有限量子博弈中的去中心化全信息学习在 $T=O( max_i d_i log d_i/ ε^2)$ 轮后达到 $ε$-近似可分离量子相关均衡。本文将此类均衡与中介量子推荐协议的通道鲁棒性关联,并给出适用于任意有限维态的局部CPTP可利用性半定规划审计,还包含基于哈随机纯态探测的探测带宽扩展,伪后悔为 $O(d^{4/3}T^{2/3}( log d)^{1/3})$。
原文摘要 · Abstract (English)
External regret certifies stability only against replacing one's behavior by a fixed alternative. In a quantum game, this misses a natural physical move: a player can apply a local completely positive trace-preserving (CPTP) map to the state it actually received or prepared. We introduce coherent swap regret as the regret benchmark against all such local CPTP deviations, and give an algorithm achieving $O(\sqrt{dT\log d})$ coherent swap regret via entropic mirror ascent on the CPTP Choi slice with a fixed-point play rule. The main result is a three-level deviation-class landscape. Replacement channels recover ordinary external regret at rate $Θ(\sqrt{T\log d})$. Unital channels, including unitary deviations and mixtures of unitaries, have zero minimax regret. Deterministic measurement-and-preparation channels already force $Ω(\sqrt{dT\log d})$ regret in the moderate-horizon regime, and this rate is also sufficient for all CPTP deviations. Thus the hardness comes from non-unital use of the recommendation register, not from quantum coherence alone. As an application, decentralized full-information learning in finite quantum games reaches an $\varepsilon$-approximate separable quantum correlated equilibrium after $T=O(\max_i d_i\log d_i/\varepsilon^2)$ rounds. We identify these equilibria with channel-proofness of mediated quantum recommendation protocols, give an SDP audit for local CPTP exploitability applicable to arbitrary finite-dimensional states, and include a probing-bandit extension with pseudo-regret $O(d^{4/3}T^{2/3}(\log d)^{1/3})$ under Haar-random pure-state probes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。