探索表征学习中新的距离度量,提升聚类与降维性能
Beyond I-Con: Exploring New Dimension of Distance Measures in Representation Learning
- 用总变差距离改进PMI算法,实现无监督聚类新最佳
- 以JSD替代标准损失,提升基于欧氏距离的对比学习效果
- 用有界f散度替换KL,降维后下游任务表现更优
信息对比(I-Con)框架揭示,超过23种表征学习方法隐式最小化数据与所学分布之间的KL散度,该散度衡量数据点间相似性。然而,基于KL的损失可能与真实目标不一致,且其不对称性和无界性会带来优化挑战。本文提出Beyond I-Con框架,系统探索替代统计散度以发现新型损失函数。关键发现:(1) 在DINO-ViT嵌入的无监督聚类任务中,将PMI算法改为使用总变差(TV)距离,达到当前最优结果;(2) 使用欧氏距离作为特征空间度量的监督对比学习中,用Jensen-Shannon散度(JSD)替代标准损失函数,显著提升性能;(3) 在降维任务中,用有界f散度替代KL散度,相比t-SNE获得更优的可视化效果和下游任务表现。结果表明,表征学习中的散度选择对优化至关重要。
原文摘要 · Abstract (English)
The Information Contrastive (I-Con) framework revealed that over 23 representation learning methods implicitly minimize KL divergence between data and learned distributions that encode similarities between data points. However, a KL-based loss may be misaligned with the true objective, and properties of KL divergence such as asymmetry and unboundedness may create optimization challenges. We present Beyond I-Con, a framework that enables systematic discovery of novel loss functions by exploring alternative statistical divergences. Key findings: (1) on unsupervised clustering of DINO-ViT embeddings, we achieve state-of-the-art results by modifying the PMI algorithm to use total variation (TV) distance; (2) supervised contrastive learning with Euclidean distance as the feature space metric is improved by replacing the standard loss function with Jenson-Shannon divergence (JSD); (3) on dimensionality reduction, we achieve superior qualitative results and better performance on downstream tasks than SNE by replacing KL with a bounded $f$-divergence. Our results highlight the importance of considering divergence choices in representation learning optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。