随机初始化网络也能通过同伴共识学习有效表征
Randomly Initialized Networks Can Learn from Peer-to-Peer Consensus

- 仅用随机初始化网络组进行同伴间一致性训练
- 下游任务性能显著优于随机基线
- 揭示自蒸馏在学习动态中的核心作用,适合研究表征学习机制者
自监督学习中,自蒸馏方法表现出色,能学习对下游任务有用的表征,甚至展现出涌现特性。但当前顶尖方法通常依赖复杂组件的集成,设计选择多为经验性且缺乏理论理解。本文探究自蒸馏在学习动态中的作用,通过训练一组随机初始化网络,移除投影器、预测器及预训练任务等常见组件,仅保留自蒸馏机制。结果表明,即使在这一极简设置下,模型仍能在下游任务中实现显著优于随机基线的表示能力。我们还分析了不同超参数对效果的影响,并探讨了该设置下模型实际学习的内容。
原文摘要 · Abstract (English)
In self-supervised learning, self-distilled methods have shown impressive performance, learning representations useful for downstream tasks and even displaying emergent properties. However, state-of-the-art methods usually rely on ensembles of complex mechanisms, with many design choices that are empirically motivated and not well understood. In this work, we explore the role of self-distillation within learning dynamics. Specifically, we isolate the effect of self-distillation by training a group of randomly initialized networks, removing all other common components such as projectors, predictors, and even pretext tasks. Our findings show that even this minimal setup can lead to learned representations with non-trivial improvements over a random baseline on downstream tasks. We also demonstrate how this effect varies with different hyperparameters and present a short analysis of what is being learned by the models under this setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。