用两次自监督替代伪标签,让聚类更稳定可靠。
Rethinking Deep Clustering Paradigms: Self-Supervision Is All You Need
- 用两轮自监督训练替代伪标签,避免特征混乱
- 在6个数据集上显著提升聚类性能
- 适合追求高稳定性的无监督学习研究者
深度聚类的最新进展得益于自监督与伪监督学习的显著进步。然而,自监督与伪监督之间的权衡会引发三大问题:联合训练导致特征随机性与特征漂移,独立训练则引发特征随机性与特征扭曲。本质上,使用伪标签会产生随机且不可靠的特征,而伪监督与自监督的结合会使面向聚类的可靠特征发生漂移。此外,从自监督转向伪监督还会扭曲潜在流形的曲率。本文针对现有深度聚类范式中的特征随机性、特征漂移和特征扭曲问题,提出一种新范式:以第二轮自监督训练取代伪监督。该策略使实例级自监督与邻域级自监督之间的过渡更平滑,减少突变;同时抑制了实例级自监督与聚类级伪监督之间的强竞争带来的漂移效应。由于没有伪监督,也规避了生成随机特征的风险。所提方法称为R-DC,专门应对深度聚类中的三大挑战。在六个数据集上的实验表明,两阶段自监督训练带来了显著性能提升。
原文摘要 · Abstract (English)
The recent advances in deep clustering have been made possible by significant progress in self-supervised and pseudo-supervised learning. However, the trade-off between self-supervision and pseudo-supervision can give rise to three primary issues. The joint training causes Feature Randomness and Feature Drift, whereas the independent training causes Feature Randomness and Feature Twist. In essence, using pseudo-labels generates random and unreliable features. The combination of pseudo-supervision and self-supervision drifts the reliable clustering-oriented features. Moreover, moving from self-supervision to pseudo-supervision can twist the curved latent manifolds. This paper addresses the limitations of existing deep clustering paradigms concerning Feature Randomness, Feature Drift, and Feature Twist. We propose a new paradigm with a new strategy that replaces pseudo-supervision with a second round of self-supervision training. The new strategy makes the transition between instance-level self-supervision and neighborhood-level self-supervision smoother and less abrupt. Moreover, it prevents the drifting effect that is caused by the strong competition between instance-level self-supervision and clustering-level pseudo-supervision. Moreover, the absence of the pseudo-supervision prevents the risk of generating random features. With this novel approach, our paper introduces a Rethinking of the Deep Clustering Paradigms, denoted by R-DC. Our model is specifically designed to address three primary challenges encountered in Deep Clustering: Feature Randomness, Feature Drift, and Feature Twist. Experimental results conducted on six datasets have shown that the two-level self-supervision training yields substantial improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。