提出适用于密集关联记忆模型的超参数迁移方法,解决层间共享权重与快速激活函数带来的挑战。
Hyperparameter Transfer for Dense Associative Memories

- 基于能量景观的动态特性,推导出小模型超参数向大规模模型迁移的理论公式
- 实验证明理论预测与实际训练结果高度一致,迁移精度接近原生调参效果
- 适合研究新型神经网络架构或需高效调参的密集关联记忆应用者
密集关联记忆(Dense Associative Memory, DenseAM)是一类具有前景的AI架构,其通过神经网络在能量景观上的时序动态实现计算。尽管前馈网络的超参数迁移方法已较成熟,但针对权重在层间及层内共享、且使用快速峰值激活函数的DenseAM场景,现有方法尚不适用。本文首次系统性地开展DenseAM超参数迁移研究,推导出小模型上调优的超参数如何准确迁移到大规模训练中的理论依据。实验验证了理论预测与实际性能之间的一致性,表明该方法可有效提升训练效率与泛化能力。
原文摘要 · Abstract (English)
Dense Associative Memory (DenseAM) is a promising family of AI architectures that is represented by a neural network performing temporal dynamics on an energy landscape. While hyperparameter transfer methods are well-studied for feed-forward networks, these methods have not been developed for settings in which weights are shared across layers and within the layer, which is common in DenseAMs. Additionally, DenseAMs utilize rapidly peaking activation functions that are rarely used in feed-forward architectures. The confluence of these aspects makes DenseAM a challenging framework for using existing methods for hyperparameter transfer. Our work initiates the development of hyperparameter transfer methods for this class of models. We derive explicit prescriptions for how the hyperparameters tuned on small models can be transferred to models trained at scale. We demonstrate excellent agreement between these theoretical findings and empirical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。