提出可保留函数的残差扩展方法,提升网络架构搜索效率
Towards a More Complete Theory of Function Preserving Transforms
- 新方法R2R在保持功能前提下整合残差连接
- 性能媲美现有方法,支持更复杂的网络结构扩展
- 适合快速训练与多样滤波器学习,适用于架构搜索
本文提出新型神经网络架构变换技术,可在不改变网络功能的前提下调整其结构。这类操作称为函数保持变换,已用于网络间知识迁移和快速架构评估,对高效架构搜索具有重要意义。我们提出的R2R方法首次将残差连接融入函数保持变换,提供理论推导并验证其性能与现有方法如Net2Net和Network Morphisms相当,从而放宽了可扩展架构的限制。通过对比分析,揭示了各类方法的差异与适用场景。实验表明,R2R能显著加快模型训练速度,并在图像分类任务中学习到更丰富的滤波器集合,优于Net2Net和Network Morphisms。
原文摘要 · Abstract (English)
In this paper, we develop novel techniques that can be used to alter the architecture of a neural network, while maintaining the function it represents. Such operations are known as function preserving transforms and have proven useful in transferring knowledge between networks to evaluate architectures quickly, thus having applications in efficient architectures searches. Our methods allow the integration of residual connections into function preserving transforms, so we call them R2R. We provide a derivation for R2R and show that it yields competitive performance with other function preserving transforms, thereby decreasing the restrictions on deep learning architectures that can be extended through function preserving transforms. We perform a comparative analysis with other function preserving transforms such as Net2Net and Network Morphisms, where we shed light on their differences and individual use cases. Finally, we show the effectiveness of R2R to train models quickly, as well as its ability to learn a more diverse set of filters on image classification tasks compared to Net2Net and Network Morphisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。