arXiv:2501.18452cs.LGcs.AI2025-01ICML被引 8

利用自监督学习的聚类特性,实现模型自我优化,提升表示质量。

Clustering Properties of Self-Supervised Learning

  • 通过聚类分析发现编码器输出具有更强更稳定的聚类特性
  • 提出ReSA方法,利用聚类特性实现自反馈训练,性能显著超越现有方法
  • 适合研究自监督学习、表征学习的科研人员参考

基于联合嵌入架构的自监督学习方法在无标签监督情况下展现出强大的语义丰富表示能力,其聚类特性尤为突出。然而,现有方法极少利用这些未被挖掘的特性进行自我改进。本文通过多种度量验证,发现编码器输出的嵌入(encoding)相比其他组件具备更优且更稳定的聚类特性。基于此,我们提出一种新型自反馈自监督学习方法——表示自分配(ReSA),利用模型自身的聚类特性实现自引导学习。在标准自监督学习基准上的大量实验表明,使用ReSA预训练的模型显著优于当前最先进的自监督方法。最后,我们分析了ReSA如何促进更好的聚类性能,证明其在细粒度与粗粒度层面均有效提升聚类效果,生成更具结构化和语义意义的表示。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) methods via joint embedding architectures have proven remarkably effective at capturing semantically rich representations with strong clustering properties, magically in the absence of label supervision. Despite this, few of them have explored leveraging these untapped properties to improve themselves. In this paper, we provide an evidence through various metrics that the encoder's output $encoding$ exhibits superior and more stable clustering properties compared to other components. Building on this insight, we propose a novel positive-feedback SSL method, termed Representation Self-Assignment (ReSA), which leverages the model's clustering properties to promote learning in a self-guided manner. Extensive experiments on standard SSL benchmarks reveal that models pretrained with ReSA outperform other state-of-the-art SSL methods by a significant margin. Finally, we analyze how ReSA facilitates better clustering properties, demonstrating that it effectively enhances clustering performance at both fine-grained and coarse-grained levels, shaping representations that are inherently more structured and semantically meaningful.

自监督学习聚类特性表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。