arXiv:2511.05462cs.LGcs.CV2025-11

将聚类方法与统计混合模型结合,提升无监督学习性能。

SiamMM: A Mixture Model Perspective on Deep Unsupervised Learning

  • 从统计混合模型视角重构聚类方法,提供理论支撑
  • 在多个自监督基准上达到当前最优表现
  • 发现聚类结果与真实标签高度相似,或可揭示标注错误

近期研究证明了基于聚类的自监督与无监督学习方法的有效性。然而,聚类应用常依赖经验性策略,最佳方法尚不明确。本文建立这些无监督聚类方法与统计学中经典混合模型之间的联系。基于该框架,我们显著改进了聚类方法,提出新型模型 SiamMM。该方法在多个自监督学习基准上取得领先性能。对学习到的聚类进行分析发现,其与未见真实标签高度一致,可能揭示了数据中标注错误的潜在实例。

原文摘要 · Abstract (English)

Recent studies have demonstrated the effectiveness of clustering-based approaches for self-supervised and unsupervised learning. However, the application of clustering is often heuristic, and the optimal methodology remains unclear. In this work, we establish connections between these unsupervised clustering methods and classical mixture models from statistics. Through this framework, we demonstrate significant enhancements to these clustering methods, leading to the development of a novel model named SiamMM. Our method attains state-of-the-art performance across various self-supervised learning benchmarks. Inspection of the learned clusters reveals a strong resemblance to unseen ground truth labels, uncovering potential instances of mislabeling.

无监督学习聚类混合模型自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。