arXiv:2605.08885cs.LG2026-05

通过结构剪枝提升原子模型精度与推理效率的平衡

Compact SO(3) Equivariant Atomistic Foundation Models via Structural Pruning

论文配图:Compact SO(3) Equivariant Atomistic Foundation Models via Structural Pruning
图 1 · 摘自论文原文
  • 按不可约表示块整体剪枝,保持SO(3)等变性
  • 剪枝后模型在Matbench上7项指标超越从零训练的小模型
  • 可兼容量化与知识蒸馏,适合高效部署原子模型

SO(3)等变图神经网络已成为原子基础模型的主流范式,通过在架构中直接嵌入旋转对称性实现高精度与数据高效。然而其高阶张量操作带来的计算开销,导致模型精度与推理效率间存在严峻权衡。本文提出一种针对SO(3)等变原子基础模型的结构剪枝方法,沿通道与阶数维度进行剪枝,每个不可约表示作为完整块保留或移除,从而保持SO(3)等变性。从大模型检查点开始剪枝,所得模型显著降低推理成本,同时精度高于独立训练的小模型。剪枝后的MACE-MP模型在Matbench Discovery排行榜上9项指标中有7项优于官方从零训练的小模型。在效率方面,压缩后的MACE-MP与MACE-OFF模型参数量减少1.5×至4×,预训练计算量减少2.5×至4×。下游任务中,微调剪枝模型相较从零训练的任务专用模型,能量与力误差分别降低70.1%和34.4%(在8个代表性数据集上)。该方法可推广至其他SO(3)等变架构(SevenNet、eSCN),并可与量化、知识蒸馏结合以进一步提升性能。

原文摘要 · Abstract (English)

SO(3) equivariant graph neural networks have become the dominant paradigm for atomistic foundation models, achieving high accuracy and data efficiency by building rotational symmetry directly into the architecture. Yet the computational cost of their higher-order tensor operations creates a tough trade-off between model accuracy and inference efficiency. In this paper, we propose a structural pruning method for SO(3) equivariant atomistic foundation models to bridge this accuracy-efficiency gap. The pruning is applied along the channel and order dimensions, with each irreducible representation kept or removed as a complete block, thereby retaining SO(3) equivariance. Starting from a large checkpoint, the pruned model substantially reduces the inference cost while retaining higher accuracy than an independently trained small model. The pruned MACE-MP model outperforms the official from-scratch trained small model on 7 of 9 metrics on the Matbench Discovery leaderboard. In terms of efficiency, compressed MACE-MP and MACE-OFF models contain 1.5$\times$ to 4$\times$ fewer parameters and require 2.5$\times$ to 4$\times$ less pre-training compute than training a small model from scratch. For downstream applications, fine-tuning the pruned model reduces energy and force errors by 70.1% and 34.4% compared to training task-specific models from scratch across eight representative downstream datasets. We demonstrate that the method generalizes to other SO(3) equivariant architectures (SevenNet, eSCN) and can be combined with quantization and knowledge distillation for further gains.

原子模型等变网络模型压缩结构剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。