arXiv:2504.04647cs.LGq-bio.QM2025-04

通过动态计算药物类别间距离,提升长尾分类中稀有类别的识别能力。

Sub-Clustering for Class Distance Recalculation in Long-Tailed Drug Classification

  • 用子聚类对比学习挖掘类别特征,动态评估类别间可分性。
  • 在多个数据集上提升尾部类别准确率,且不损害头部类别性能。
  • 适合关注药物分类、长尾分布问题的研究者使用。

现实世界中长尾数据分布普遍,导致模型难以有效学习和分类尾部类别。然而我们发现,在药物化学领域,部分尾部类别因独特的分子结构特征,在训练过程中表现出更高的可辨识性,这与传统认为尾部类别普遍难识别的观点显著不同。现有不平衡学习方法如重采样和代价敏感加权,过度依赖样本数量先验,使模型过度关注尾部类别而牺牲头部类别表现。为此,我们提出一种新方法,突破传统基于样本数量的静态评估范式,建立基于类别间特征距离的动态类间可分性度量。具体而言,采用子聚类对比学习充分学习每个类别的嵌入特征,并动态计算类别嵌入间的距离,捕捉不同类别样本在特征空间中的相对位置演化,从而重新平衡分类损失权重。我们在多个现有的长尾药物数据集上进行实验,结果表明该方法在不降低主导类别性能的前提下,显著提升了尾部类别的准确率,达到具有竞争力的性能。

原文摘要 · Abstract (English)

In the real world, long-tailed data distributions are prevalent, making it challenging for models to effectively learn and classify tail classes. However, we discover that in the field of drug chemistry, certain tail classes exhibit higher identifiability during training due to their unique molecular structural features, a finding that significantly contrasts with the conventional understanding that tail classes are generally difficult to identify. Existing imbalance learning methods, such as resampling and cost-sensitive reweighting, overly rely on sample quantity priors, causing models to excessively focus on tail classes at the expense of head class performance. To address this issue, we propose a novel method that breaks away from the traditional static evaluation paradigm based on sample size. Instead, we establish a dynamical inter-class separability metric using feature distances between different classes. Specifically, we employ a sub-clustering contrastive learning approach to thoroughly learn the embedding features of each class, and we dynamically compute the distances between class embeddings to capture the relative positional evolution of samples from different classes in the feature space, thereby rebalancing the weights of the classification loss function. We conducted experiments on multiple existing long-tailed drug datasets and achieved competitive results by improving the accuracy of tail classes without compromising the performance of dominant classes.

长尾分类药物化学对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。