发现量子机器学习中表示与泛化能力的普适标度律
Universal scaling laws in quantum-probabilistic machine learning by tensor network towards interpreting representation and generalization powers
- 用张量网络生成模型揭示负对数似然随特征数线性增长
- 训练后出现负二次项修正,体现泛化能力提升
- 标度系数对样本数和量子通道数呈对数关系,可判过参数
解释机器学习中的表征与泛化能力是长期难题。本文揭示了量子概率机器学习中普适标度律的涌现。以矩阵乘积态形式的生成张量网络(GTN)为例,未训练的随机张量网络状态下,负对数似然(NLL)L 随特征数 M 线性增长,即 $L \simeq k M + const$,源于‘正交灾难’——随着 M 增大,量子多体态趋于指数正交。训练过程中,信息获取抑制了线性增长,引入负二次修正项,形成 $L \simeq βM - αM^2 + const$。该二次项在测试集(训练集)上的出现,可视为泛化(表征)能力的证据。通过比较 α 值在训练与测试集间的偏差,可识别过参数化现象。研究还揭示了量子特征映射正交性与量子概率解释、以及表征与泛化能力之间的关联。该工作为建立基于量子概率框架的白盒机器学习提供了关键步骤。
原文摘要 · Abstract (English)
Interpreting the representation and generalization powers has been a long-standing issue in the field of machine learning (ML) and artificial intelligence. This work contributes to uncovering the emergence of universal scaling laws in quantum-probabilistic ML. We take the generative tensor network (GTN) in the form of a matrix product state as an example and show that with an untrained GTN (such as a random TN state), the negative logarithmic likelihood (NLL) $L$ generally increases linearly with the number of features $M$, i.e., $L \simeq k M + const$. This is a consequence of the so-called ``catastrophe of orthogonality,'' which states that quantum many-body states tend to become exponentially orthogonal to each other as $M$ increases. We reveal that while gaining information through training, the linear scaling law is suppressed by a negative quadratic correction, leading to $L \simeq βM - αM^2 + const$. The scaling coefficients exhibit logarithmic relationships with the number of training samples and the number of quantum channels $χ$. The emergence of the quadratic correction term in NLL for the testing (training) set can be regarded as evidence of the generalization (representation) power of GTN. Over-parameterization can be identified by the deviation in the values of $α$ between training and testing sets while increasing $χ$. We further investigate how orthogonality in the quantum feature map relates to the satisfaction of quantum probabilistic interpretation, as well as to the representation and generalization powers of GTN. The unveiling of universal scaling laws in quantum-probabilistic ML would be a valuable step toward establishing a white-box ML scheme interpreted within the quantum probabilistic framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。