研究大模型在量子化学中的缩放规律,揭示其性能与资源的非线性关系。
Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry
- 基于Transformer的神经量子态模型性能随问题规模变化有可预测缩放规律。
- 绝对误差和V-score指标显示,模型大小与训练时间呈非线性依赖关系。
- 结果对优化量子化学模拟的算力分配具有指导意义,适合量子计算研究者。
缩放规律被用于描述大语言模型(LLM)性能如何随模型规模、训练数据量或计算资源的变化而变化。鉴于神经量子态(NQS)越来越多地采用基于LLM的组件,我们旨在理解NQS的缩放规律,从而揭示其可扩展性及性能-资源权衡的最优策略。具体而言,我们识别出适用于第二量化量子化学应用中基于Transformer的NQS的缩放规律,该规律能预测其性能(以绝对误差和V-score衡量)随问题规模的变化。通过对获得的参数曲线进行类似的计算受限优化,发现模型规模与训练时间的关系高度依赖于损失函数和波函数形式,且不遵循语言模型中观察到的近似线性关系。
原文摘要 · Abstract (English)
Scaling laws have been used to describe how large language model (LLM) performance scales with model size, training data size, or amount of computational resources. Motivated by the fact that neural quantum states (NQS) has increasingly adopted LLM-based components, we seek to understand NQS scaling laws, thereby shedding light on the scalability and optimal performance--resource trade-offs of NQS ansatze. In particular, we identify scaling laws that predict the performance, as measured by absolute error and V-score, for transformer-based NQS as a function of problem size in second-quantized quantum chemistry applications. By performing analogous compute-constrained optimization of the obtained parametric curves, we find that the relationship between model size and training time is highly dependent on loss metric and ansatz, and does not follow the approximately linear relationship found for language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。