等变神经网络在大规模训练中更高效,但数据增强可弥补差距。
Does equivariance matter at scale?
- 用等变结构设计模型提升数据效率
- 等变模型在各算力下均表现更优,遵循幂律增长
- 算力分配策略对等变与非等变模型不同
在大规模数据和充足算力下,为每个问题的结构与对称性设计神经架构是否有益?还是应让模型从数据中自行学习?我们通过实验研究了等变与非等变网络在算力和训练样本规模下的表现。以刚体相互作用为基准任务,使用通用Transformer架构,系统调整模型规模、训练步数和数据集大小。结果表明:第一,等变性提升数据效率,但通过数据增强训练非等变模型,足够多的训练轮次可缩小差距;第二,算力扩展遵循幂律,等变模型在每项测试算力预算下均优于非等变模型;第三,等变与非等变模型的算力分配最优策略不同。
原文摘要 · Abstract (English)
Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。