arXiv:2606.22984cond-mat.dis-nncond-mat.stat-mech2026-06被引 2

用物理启发的稀疏注意力模型,实现单卡超大规模自旋玻璃模拟。

Scalable Physics-Inspired Transformers for Spin Glasses

论文配图:Scalable Physics-Inspired Transformers for Spin Glasses
图 1 · 摘自论文原文
  • 设计可解释的稀疏注意力与自旋特化位置编码
  • 单卡模拟规模达以往100倍,速度提升百倍以上
  • 适用于各类复杂自旋玻璃模型,尤其在高温/低温区有效

在挫败性自旋玻璃中高效采样玻尔兹曼分布是统计力学与组合优化的核心问题。尽管机器学习方法取得进展,仍面临两大挑战:变分模型规模扩大时性能未如大语言模型般单调提升,缺乏理论理解;且大规模系统计算成本高,难超经典采样方法。本文提出一种物理启发的Transformer,采用可解释的稀疏注意力机制与自旋特化位置嵌入,并结合FlashAttention实现并行祖先采样,相较传统变分自回归网络提速达两个数量级,可在单张GPU上完成前所未有的大规模自旋玻璃模拟。该方法能精确解析谢林顿-柯克帕特里克及二维、三维爱德华-安德森模型的全概率分布、自由能与重叠统计,覆盖不同温度区间,突破现有机器学习方法在特定温度下的瓶颈。该框架为挫败性自旋玻璃系统建立了可扩展的模拟范式。

原文摘要 · Abstract (English)

Efficient sampling of the Boltzmann distribution in frustrated spin glasses is central to statistical mechanics and combinatorial optimization. Despite advances in machine-learning-based approaches, two issues persist: limited understanding of why variational models fail to benefit from increased scale, unlike the monotonic scaling law of large language models; and high computational cost on large systems that negates advantages over classical sampling methods. Here, we develop a physics-inspired transformer with interpretable sparse attention and spin-tailored positional embeddings to address these challenges. By further leveraging FlashAttention for parallel ancestral sampling, it achieves up to two orders of magnitude speedup over vanilla variational autoregressive networks, enabling neural-network simulations of spin-glass systems to unprecedented sizes on a single GPU. It can resolve full probability distributions, free energies, and overlap statistics across temperatures, for Sherrington-Kirkpatrick and 2D or 3D Edwards-Anderson models, where existing machine-learning methods encounter limitations at certain temperatures. This framework thus establishes a scalable paradigm for frustrated spin-glass systems.

自旋玻璃Transformer物理模型高效采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。