用密度信息设计学习策略,提升图像聚类的鲁棒性与速度
Deep Image Clustering Based on Curriculum Learning and Density Information

- 基于数据密度设计课程学习策略,优化模型训练节奏
- 用聚类核心代替中心点指导聚类分配,减少误差累积
- 在多个数据集上表现更优,适应不同规模和场景
图像聚类是多媒体分析与知识发现中的关键技术。近年来,深度聚类方法(DC)因其能联合进行特征学习与聚类分配,在图像数据上超越了传统方法。然而,现有方法很少关注模型学习策略对复杂图像聚类性能和鲁棒性的提升作用。此外,大多数方法仅依赖点到聚类中心的距离进行潜在表示划分,导致迭代过程中误差累积。本文提出一种鲁棒的图像聚类方法(IDCL),首次将密度信息引入聚类训练策略中。具体而言,我们基于输入数据的密度信息设计了一种课程学习方案,实现更合理的学习节奏;同时,采用聚类核心而非单个聚类中心来指导聚类分配。在基准数据集上的大量实验表明,所提方法在鲁棒性、快速收敛及数据规模、聚类数和图像上下文方面的灵活性上均优于当前最优方法。
原文摘要 · Abstract (English)
Image clustering is one of the crucial techniques in multimedia analytics and knowledge discovery. Recently, the Deep clustering method (DC), characterized by its ability to perform feature learning and cluster assignment jointly, surpasses the performance of traditional ones on image data. However, existing methods rarely consider the role of model learning strategies in improving the robustness and performance of clustering complex image data. Furthermore, most approaches rely solely on point-to-point distances to cluster centers for partitioning the latent representations, resulting in error accumulation throughout the iterative process. In this paper, we propose a robust image clustering method (IDCL) which, to our knowledge for the first time, introduces a model training strategy using density information into image clustering. Specifically, we design a curriculum learning scheme grounded in the density information of input data, with a more reasonable learning pace. Moreover, we employ the density core rather than the individual cluster center to guide the cluster assignment. Finally, extensive comparisons with state-of-the-art clustering approaches on benchmark datasets demonstrate the superiority of the proposed method, including robustness, rapid convergence, and flexibility in terms of data scale, number of clusters, and image context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。