arXiv:2510.09768cs.LGcs.AI2025-10被引 6

对称性让神经势能模型在更大规模下表现更好,且数据、参数、算力需协同增长。

Scaling Laws and Symmetry, Evidence from Neural Force Fields

  • 利用任务对称性设计的等变架构更易扩展。
  • 等变模型在数据、参数、算力上均呈现幂律缩放,指数由架构决定。
  • 高阶表示的等变模型缩放性能更优,适合大规模分子模拟研究者。

我们针对学习原子间势能这一几何任务进行了实证研究,发现即使在更大规模下,对称性依然至关重要;观察到数据、参数和算力之间存在明确的幂律缩放关系,且缩放指数依赖于架构。特别地,利用任务对称性的等变架构比非等变模型具有更好的扩展性。此外,在等变架构中,高阶表征带来更优的缩放指数。分析还表明,为实现计算最优训练,数据量与模型规模应同步增长,与架构无关。总体而言,这些结果表明,不应依赖模型自行发现对称性等基本归纳偏置,尤其在规模化时,因为它们会改变任务的本质难度及其缩放规律。

原文摘要 · Abstract (English)

We present an empirical study in the geometric task of learning interatomic potentials, which shows equivariance matters even more at larger scales; we show a clear power-law scaling behaviour with respect to data, parameters and compute with ``architecture-dependent exponents''. In particular, we observe that equivariant architectures, which leverage task symmetry, scale better than non-equivariant models. Moreover, among equivariant architectures, higher-order representations translate to better scaling exponents. Our analysis also suggests that for compute-optimal training, the data and model sizes should scale in tandem regardless of the architecture. At a high level, these results suggest that, contrary to common belief, we should not leave it to the model to discover fundamental inductive biases such as symmetry, especially as we scale, because they change the inherent difficulty of the task and its scaling laws.

神经势能对称性缩放定律分子模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。