arXiv:2601.03237cs.LGeess.IV2026-01被引 1

解决数据不平衡下的聚类误差问题,提升小样本类别识别准确率。

PET-TURTLE: Deep Unsupervised Support Vector Machines for Imbalanced Data Clusters

  • 引入幂律先验优化损失函数,适应不均衡数据分布。
  • 通过稀疏逻辑值缩小搜索空间,提升平衡数据集精度。
  • 在真实与合成数据上均有效降低少数类过预测问题。

基础视觉、音频和语言模型通过其隐式表征实现下游任务的零样本性能。近年来,基于深度学习方法的无监督数据结构学习受到关注。TURTLE 是一种先进的深度聚类算法,通过交替更新标签和超平面,最大化超平面间隔,类似支持向量机(SVM)。然而,TURTLE 假设簇是平衡的;当数据不平衡时,会产生非理想超平面,导致更高聚类误差。我们提出 PET-TURTLE,通过幂律先验推广损失函数以处理不均衡数据分布。此外,通过在标记过程中引入稀疏逻辑值,PET-TURTLE 优化更简单的搜索空间,从而提升平衡数据集的准确性。在合成与真实数据上的实验表明,PET-TURTLE 改善了不均衡数据源的聚类准确率,防止对少数簇的过预测,并提升了整体聚类性能。

原文摘要 · Abstract (English)

Foundation vision, audio, and language models enable zero-shot performance on downstream tasks via their latent representations. Recently, unsupervised learning of data group structure with deep learning methods has gained popularity. TURTLE, a state of the art deep clustering algorithm, uncovers data labeling without supervision by alternating label and hyperplane updates, maximizing the hyperplane margin, in a similar fashion to support vector machines (SVMs). However, TURTLE assumes clusters are balanced; when data is imbalanced, it yields non-ideal hyperplanes that cause higher clustering error. We propose PET-TURTLE, which generalizes the cost function to handle imbalanced data distributions by a power law prior. Additionally, by introducing sparse logits in the labeling process, PET-TURTLE optimizes a simpler search space that in turn improves accuracy for balanced datasets. Experiments on synthetic and real data show that PET-TURTLE improves accuracy for imbalanced sources, prevents over-prediction of minority clusters, and enhances overall clustering.

聚类不平衡数据深度学习无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。