无需计算距离,用神经网络自动聚类数据。
Cohort Organized Learning: Clustering Through Agreement

- 通过期望最大化推导梯度,训练神经网络实现聚类。
- 可处理向量与图像数据,支持任意兼容输入。
- 适合无标签数据聚类,尤其适合复杂结构数据。
本文介绍了一种名为群体组织学习(Cohort Organized Learning, CoOL)的方法,用于在不显式计算距离或相似性的情况下对数据进行聚类。文中详细描述了CoOL的原理,推导了基于期望最大化算法的梯度以训练神经网络,展示了训练过程中的收敛监控方法以及训练后聚类结果的评估方式,并讨论了一系列应用实例与使用场景。同时,文章也探讨了CoOL的局限性及在相关任务中的未来发展方向。由于CoOL利用神经网络估计聚类结构,因此可适用于任何可适配的数据类型,我们已在向量数据和图像数据上进行了验证。
原文摘要 · Abstract (English)
In this article we describe Cohort Organized Learning (CoOL), a method for clustering data without explicit distance or similarity computations. Herein, we will describe CoOL, derive the gradients determined by expectation maximization to train the networks, show how to monitor convergence during training and evaluate the clusters after training, and discuss a series of examples and use cases. We also discuss CoOL's limitations and future prospects on related tasks. Because CoOL uses neural networks to estimate the clusters, it can be used to cluster any data that can be made compatible and we illustrate this on vector data and images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。