用新型神经网络提升量子化学计算效率,速度更快且不损失精度。
Retentive Neural Quantum States: Efficient Ansätze for Ab Initio Quantum Chemistry
- 采用回溯网络(RetNet)替代Transformer,实现训练并行、推理递归。
- 在问题规模大于模型规模时,计算时间比Transformer快2倍以上。
- 通过变分神经退火策略弥补表达能力差距,适合大规模量子模拟研究者。
神经网络量子态(NQS)作为量子启发的深度学习在变分蒙特卡洛方法中的重要应用,为求解量子体系基态提供了有竞争力的方案。近期将自回归模型(尤其是Transformer)引入作为变分波函数,显著提升了表达能力,但其时间复杂度随序列长度增长而急剧上升。本文探索使用回溯网络(RetNet)作为求解原子间无参数量子化学基态问题的波函数候选。与Transformer不同,RetNet在训练阶段并行处理数据,在推理阶段递归运行,有效缓解了时间复杂度瓶颈。我们给出了RetNet的简单计算成本估算,并与Transformer进行直接对比,确立了一个明确的临界比例:当问题规模超过模型规模时,RetNet的时间复杂度优势显现。尽管表达能力略低于Transformer,我们通过利用模型自回归结构的变分神经退火策略成功弥补了这一差距。结果表明,RetNet可在不牺牲精度的前提下显著改善NQS的计算效率。进一步分析显示,神经退火的改进效果不仅限于RetNet,可推广至通用自回归型NQS,具有广泛适用性。
原文摘要 · Abstract (English)
Neural-network quantum states (NQS) has emerged as a powerful application of quantum-inspired deep learning for variational Monte Carlo methods, offering a competitive alternative to existing techniques for identifying ground states of quantum problems. A significant advancement toward improving the practical scalability of NQS has been the incorporation of autoregressive models, most recently transformers, as variational ansatze. Transformers learn sequence information with greater expressiveness than recurrent models, but at the cost of increased time complexity with respect to sequence length. We explore the use of the retentive network (RetNet), a recurrent alternative to transformers, as an ansatz for solving electronic ground state problems in $\textit{ab initio}$ quantum chemistry. Unlike transformers, RetNets overcome this time complexity bottleneck by processing data in parallel during training, and recurrently during inference. We give a simple computational cost estimate of the RetNet and directly compare it with similar estimates for transformers, establishing a clear threshold ratio of problem-to-model size past which the RetNet's time complexity outperforms that of the transformer. Though this efficiency can comes at the expense of decreased expressiveness relative to the transformer, we overcome this gap through training strategies that leverage the autoregressive structure of the model -- namely, variational neural annealing. Our findings support the RetNet as a means of improving the time complexity of NQS without sacrificing accuracy. We provide further evidence that the ablative improvements of neural annealing extend beyond the RetNet architecture, suggesting it would serve as an effective general training strategy for autoregressive NQS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。