研究3D旋转等变性的训练速度与效果,发现模型很快就能学会对称性。
Training Dynamics of Learning 3D-Rotational Equivariance
- 用可量化误差衡量模型对称性学习程度,基于凸损失函数设计
- 在分子任务中,1000~1万步内等变误差降至≤2%的测试损失
- 等变性学习比主任务更容易,适合提升训练效率的研究者
尽管数据增强广泛用于训练对称性无关模型,但其学习对称性的速度和效果仍不明确。本文推导出一种合理的等变性误差度量方法,针对凸损失函数,计算总损失中因对称性学习不完善所占比例。我们聚焦于高维分子任务中的3D旋转等变性(包括流匹配、力场预测、去噪体素),发现模型在1000至10000个训练步内即可将等变性误差降至≤2%的保留损失,该结果对模型和数据集规模均鲁棒。这是因为3D旋转等变性学习本身是更简单的任务,具有更平滑且条件更好的损失景观。在整个训练过程中,非等变模型的损失惩罚较小,因此在单位GPU时长下可能获得更低的测试损失,除非缩小等变模型的‘效率差距’。我们还通过实验与理论分析了相对等变性误差、学习梯度与模型参数间的关系。
原文摘要 · Abstract (English)
While data augmentation is widely used to train symmetry-agnostic models, it remains unclear how quickly and effectively they learn to respect symmetries. We investigate this by deriving a principled measure of equivariance error that, for convex losses, calculates the percent of total loss attributable to imperfections in learned symmetry. We focus our empirical investigation to 3D-rotation equivariance on high-dimensional molecular tasks (flow matching, force field prediction, denoising voxels) and find that models reduce equivariance error quickly to $\leq$2\% held-out loss within 1k-10k training steps, a result robust to model and dataset size. This happens because learning 3D-rotational equivariance is an easier learning task, with a smoother and better-conditioned loss landscape, than the main prediction task. For 3D rotations, the loss penalty for non-equivariant models is small throughout training, so they may achieve lower test loss than equivariant models per GPU-hour unless the equivariant ``efficiency gap'' is narrowed. We also experimentally and theoretically investigate the relationships between relative equivariance error, learning gradients, and model parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。