用强化学习视角优化量子态模型,提升训练稳定性和效率。
One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective
- 将波函数优化转化为优势策略梯度问题,采用信任区域方法
- 在多种自旋系统上比Adam等方法收敛更快、更稳定
- 可训练15亿参数模型,规模远超以往工作
神经量子态(NQS)为近似量子多体波函数提供了灵活且可扩展的框架。其中自回归模型因其能从波恩分布中精确独立采样而备受青睐,避免了马尔可夫链方法中的自相关和混合问题。然而其优化仍相对不足:Adam虽可扩展但忽略函数空间几何,随机重构虽理论严谨但计算昂贵且数值不稳定。为此,我们提出将变分能量最小化视为波恩分布上的优势策略梯度问题,从而启发使用信任区域优化。我们引入近端波函数优化(PWO),一种在振幅通道剪裁概率比变化、相位通道限制相位增量的可信区域算法。PWO避免显式矩阵求逆,复用样本进行多轮更新,结合一阶优化的可扩展性与理论保证。在伊辛及受挫的$J_1$-$J_2$一维和二维自旋系统中,PWO在稳定性和实际运行时间上均优于Adam、minSR和SPRING。最后,我们对一个15亿参数的RWKV-7模型进行了微调,展示了规模超过以往三数量级的NQS优化。
原文摘要 · Abstract (English)
Neural quantum states (NQS) provide a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, autoregressive models are especially attractive because they enable exact, independent sampling from the Born distribution, avoiding the autocorrelation and mixing issues of Markov chain methods. Yet their optimization remains comparatively underexplored: Adam is a scalable method but ignores function space geometry, while stochastic reconfiguration is principled but costly and numerically fragile in large models. To address this gap, we show that variational energy minimization can be viewed as an advantage policy-gradient problem over the Born distribution, motivating trust-region optimization for NQS training. We introduce Proximal Wavefunction Optimization (PWO), a principled trust-region algorithm that clips probability-ratio changes in the amplitude channel and phase increments in the phase channel. PWO avoids explicit matrix inversion, reuses samples across multiple updates, and combines the scalability of first-order optimization with theoretical guarantees. Across Ising and frustrated $J_1$-$J_2$ one- and two-dimensional spin systems, PWO improves stability and wall-clock convergence over Adam, minSR, and SPRING. Finally, we fine-tune a $1.5$B-parameter RWKV-7 model, demonstrating NQS optimization at a scale over three orders of magnitude beyond prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。