DINO类自监督方法存在原型冗余问题,影响细粒度表征学习。
On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods
- 通过鼓励使用多样原型缓解部分原型坍缩
- 在长尾细粒度数据集上显著提升聚类细粒度
- 适合需要精细表征的自监督预训练场景
自监督学习常将表示建模为聚类或混合模型。同时学习紧凑表示并拟合混合模型易导致表示坍缩。主流策略是正则化数据点在聚类间的分布以避免完全坍缩。然而我们发现,DINO系列方法仍存在部分原型坍缩问题,导致原型冗余。这些冗余为模型提供捷径,使其仅需满足预定先验的隐层类别分布即可。通过鼓励模型使用多样化原型,可有效缓解该问题。充分使用原型使方法能学习更细粒度的聚类,获得更具信息量的表示。我们在长尾细粒度数据集上的预训练实验验证了这一优势。
原文摘要 · Abstract (English)
A prominent self-supervised learning paradigm is to model the representations as clusters, or more generally as a mixture model. Learning to map the data samples to compact representations and fitting the mixture model simultaneously leads to the representation collapse problem. Regularizing the distribution of data points over the clusters is the prevalent strategy to avoid this issue. While this is sufficient to prevent full representation collapse, we show that a partial prototype collapse problem still exists in the DINO family of methods, that leads to significant redundancies in the prototypes. Such prototype redundancies serve as shortcuts for the method to achieve a marginal latent class distribution that matches the prescribed prior. We show that by encouraging the model to use diverse prototypes, the partial prototype collapse can be mitigated. Effective utilization of the prototypes enables the methods to learn more fine-grained clusters, encouraging more informative representations. We demonstrate that this is especially beneficial when pre-training on a long-tailed fine-grained dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。