arXiv:2502.09252cs.LG2025-02ICML被引 9

揭示自监督学习中嵌入范数对收敛与置信度的关键作用

On the Importance of Embedding Norms in Self-Supervised Learning

论文配图:On the Importance of Embedding Norms in Self-Supervised Learning
图 1 · 摘自论文原文
  • 嵌入范数影响模型收敛速度,越小越快
  • 嵌入范数反映网络置信度,值小对应异常样本
  • 调控范数可显著加速训练,适合优化研究者

自监督学习(SSL)无需标注即可训练数据表示,已成为机器学习重要范式。多数方法使用嵌入向量的余弦相似度,隐含将数据嵌入超球面。尽管看似嵌入范数无用,近期研究指出其与网络收敛和置信度相关。本文通过理论分析、仿真与实验系统揭示:嵌入范数(i)决定SSL收敛速率,(ii)编码网络置信度,范数越小对应越意外的样本。调控嵌入范数可显著影响收敛速度。结果表明,嵌入范数是理解与优化网络行为的核心要素。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) allows training data representations without a supervised signal and has become an important paradigm in machine learning. Most SSL methods employ the cosine similarity between embedding vectors and hence effectively embed data on a hypersphere. While this seemingly implies that embedding norms cannot play any role in SSL, a few recent works have suggested that embedding norms have properties related to network convergence and confidence. In this paper, we resolve this apparent contradiction and systematically establish the embedding norm's role in SSL training. Using theoretical analysis, simulations, and experiments, we show that embedding norms (i) govern SSL convergence rates and (ii) encode network confidence, with smaller norms corresponding to unexpected samples. Additionally, we show that manipulating embedding norms can have large effects on convergence speed. Our findings demonstrate that SSL embedding norms are integral to understanding and optimizing network behavior.

自监督学习嵌入范数收敛性置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。