arXiv:2607.06230quant-phcs.LG2026-07

量子电路的泛化能力由纠缠程度决定,而非参数数量。

Entanglement as a Structural Complexity Axis: A PAC-Bayesian View of Generalization in Quantum Policies and Value Functions

  • 用费舍尔几何有效维度衡量量子策略复杂度
  • 纠缠越强,训练测试差距越大,参数量无法预测
  • 适用于量子强化学习、分类与价值函数设计

参数化量子电路(PQCs)在量子强化学习中广泛用作策略和价值函数,但其泛化机制尚不明确。本文基于PAC-Bayesian理论提出:泛化能力不取决于参数总数,而由电路诱导的费舍尔几何有效维度决定,该维度受纠缠影响而增大。在固定可训练旋转数、仅改变纠缠程度的控制实验中,费舍尔有效维度更大的电路表现出更显著的训练-测试差距,而参数数量预测力较弱。所提界主要作为排序依据,能正确区分同参数量电路,传统参数计数边界无法实现。该机制在监督分类、量子上下文赌博及价值函数泛化中均成立:同等参数下,纠缠电路泛化性能更差,且随样本量增加差距缩小。最强证据来自低方差决策模型(单观测分类器、价值头、一步策略)。端到端多步策略学习中,纠缠效应仍显著但高回报方差导致排序部分未完全解析。偏相关分析显示,费舍尔有效维度能屏蔽纠缠模式影响,且在控制训练准确率、读出方式与优化器后,排除了主要优化偏差。该效应在真实噪声环境下的IBM Heron量子处理器上依然存在。整体表明,应将量子策略设计聚焦于纠缠与泛化的权衡,而非单一表达能力。

原文摘要 · Abstract (English)

Parameterized quantum circuits (PQCs) are increasingly used as policies and value functions in quantum reinforcement learning, yet it remains unclear when and why quantum policies generalize. We give a PAC-Bayesian account in which generalization is governed not by the raw number of circuit parameters, but by the effective dimension of the Fisher geometry induced by the circuit. This quantity is inflated by entanglement, making entangling connectivity an independent axis of complexity.In controlled experiments that fix the number of trainable rotations and vary only entanglement, we find that circuits with larger Fisher effective dimension exhibit larger train-test gaps, while parameter count is a weak predictor. The resulting bound acts primarily as a ranking certificate: it correctly orders circuits with identical parameter count, which parameter-counting bounds cannot do. We validate this mechanism across supervised classification, quantum contextual bandits, and value-function generalization, where entangled circuits consistently generalize worse than non-entangled circuits of equal parameter count, with gaps shrinking as sample size increases.Our strongest evidence comes from low-variance decision models, including single-observable classifiers, value heads, and one-step policies. In end-to-end multi-step policy learning, entanglement effects remain statistically significant but high return variance leaves the full ordering only partially resolved. Partial-correlation analysis shows that Fisher effective dimension screens off entangling pattern, and controls for training accuracy, readout, and optimizer rule out major optimization confounders. The effect also persists on an IBM Heron quantum processor under real noise. Overall, our results reframe quantum policy design around an entanglement--generalization trade-off rather than expressivity alone.

量子机器学习泛化分析纠缠强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。