arXiv:2608.01434cs.LGcs.AI2026-08

将权重分布约束转化为学习的几何结构,提升模型泛化与训练稳定性。

Statistical Mechanics of Learning on Product Wasserstein Manifolds

论文配图:Statistical Mechanics of Learning on Product Wasserstein Manifolds
图 1 · 摘自论文原文
  • 把权重分布当作流形上的几何先验,构建梯度流学习框架。
  • 在图像分类与量子电路中,泛化能力提升,训练更稳定,梯度消失减弱。
  • 适合研究模型几何、量子机器学习或需引入先验知识的场景。

传统统计学习力学将权重分布约束视为对解空间的限制,导致模型容量下降。本文提出相反思路:将预设权重分布视为学习自然发生的内在几何结构。将深度神经网络与变分量子电路统一建模为多个Wasserstein流形乘积上的梯度流——每层对应一个经典Wasserstein空间,量子电路参数对应一个量子Wasserstein空间。在此几何下,容量降低表现为约束流形自身的度量结构。我们发展了深度网络的层次均场描述,通过一阶量子Wasserstein距离拓展至量子设置,并提出两种算法:Hierarchical DisCo-SGD与Quantum DisCo,沿乘积流形近似测地线演化。在教师-学生问题、标准图像分类任务及小型变分量子分类器上的实验表明,尊重分布几何可提升泛化性能、稳定训练过程,并缓解退化平原问题,优于无约束与仅基于范数的基线。该方法首次将结构约束重构为几何先验,为融合生物、谱或硬件衍生的分布信息提供新路径,适用于经典与量子学习系统。

原文摘要 · Abstract (English)

Normally the statistical mechanics of learning treats constraints on weight distributions as restrictions that shrink the space of possible solutions. Therefore, it reduces model capacity. In this paper we would like to take a contrary approach, which, however, is based on the earlier work on distribution-constrained perceptrons. Rather than treating a prescribed weight distribution as a mere restriction, we propose that it defines the intrinsic geometry upon which learning naturally unfolds. We formulate both deep neural networks and variational quantum circuits as gradient flows on a product of Wasserstein manifolds -- one classical Wasserstein space for each layer and one quantum Wasserstein space for the circuit parameters. Within this geometry, the capacity reduction, which was previously associated with distributional constraints, appears as the metric structure of the constraint manifold itself. We develop a hierarchical mean-field description for deep networks, extend the framework to the quantum setting using the quantum Wasserstein distance of order 1, and introduce two such practical algorithms, Hierarchical DisCo-SGD and Quantum DisCo, that follow approximate geodesics on the manifold of the product itself. Experiments on teacher-student problems, standard image classification tasks, and small variational quantum classifiers show that respecting these distributional geometries improves generalization, stabilizes training, and reduces the severity of barren plateaus compared with unconstrained and purely norm-based baselines. This approach firstly reframes structural constraints as geometric priors and suggests a route for incorporating biological, spectral, or hardware-derived distributional information into both learning systems, viz., classical and quantum learning.

机器学习几何学习量子机器学习分布约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。