arXiv:2412.18838cs.CVcs.LG2024-12被引 3

利用扩散模型反推文本条件,实现细粒度图像聚类

DiFiC: Your Diffusion Model Holds the Secret to Fine-Grained Clustering

  • 通过反推生成图像的文本条件获取更精确语义
  • 在四个数据集上优于当前最优聚类方法
  • 适合对细粒度分类与生成模型结合感兴趣的学者

细粒度聚类是一项实用但具有挑战性的任务,核心在于捕捉不同类别间细微差异。这些细微差异易受数据增强干扰或被冗余信息淹没,导致现有聚类方法性能显著下降。本文提出DiFiC,一种基于条件扩散模型的细粒度聚类方法。不同于聚焦于从图像中提取判别特征的现有工作,DiFiC转而推断用于图像生成的文本条件。为获得更精准、利于聚类的物体语义,DiFiC进一步对扩散目标进行正则化,并利用邻域相似性引导蒸馏过程。大量实验表明,DiFiC在四个细粒度图像聚类基准上均超越当前最先进的判别与生成式聚类方法。我们希望DiFiC的成功能激励未来研究挖掘扩散模型在生成之外任务中的潜力。代码将公开。

原文摘要 · Abstract (English)

Fine-grained clustering is a practical yet challenging task, whose essence lies in capturing the subtle differences between instances of different classes. Such subtle differences can be easily disrupted by data augmentation or be overwhelmed by redundant information in data, leading to significant performance degradation for existing clustering methods. In this work, we introduce DiFiC a fine-grained clustering method building upon the conditional diffusion model. Distinct from existing works that focus on extracting discriminative features from images, DiFiC resorts to deducing the textual conditions used for image generation. To distill more precise and clustering-favorable object semantics, DiFiC further regularizes the diffusion target and guides the distillation process utilizing neighborhood similarity. Extensive experiments demonstrate that DiFiC outperforms both state-of-the-art discriminative and generative clustering methods on four fine-grained image clustering benchmarks. We hope the success of DiFiC will inspire future research to unlock the potential of diffusion models in tasks beyond generation. The code will be released.

细粒度聚类扩散模型生成式学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。