arXiv:2503.07968cs.CL2025-03被引 4

通过共现重排序提升长尾多标签分类准确率

LabelCoRank: Revolutionizing Long Tail Multi-Label Classification with Co-Occurrence Reranking

  • 基于标签共现关系设计双阶段重排序机制
  • 在MAG-CS等数据集上显著提升长尾标签分类效果
  • 适合关注标签依赖关系的多标签分类研究者

尽管预训练大模型在语义表示方面取得进展,但多标签文本分类中的长尾问题仍难解决。现有方法多关注文本语义而忽视标签间关系。本文提出LabelCoRank,受排序思想启发,利用标签共现关系对初始分类结果进行双阶段重排序。第一阶段基于初始结果生成初步排名;第二阶段使用标签共现矩阵进一步重排,提升最终分类的准确性和相关性。通过将重排序后的标签表示作为额外文本特征,有效缓解了多标签分类中的长尾问题。在MAG-CS、PubMed和AAPD等主流数据集上的实验验证了其有效性与鲁棒性。

原文摘要 · Abstract (English)

Motivation: Despite recent advancements in semantic representation driven by pre-trained and large-scale language models, addressing long tail challenges in multi-label text classification remains a significant issue. Long tail challenges have persistently posed difficulties in accurately classifying less frequent labels. Current approaches often focus on improving text semantics while neglecting the crucial role of label relationships. Results: This paper introduces LabelCoRank, a novel approach inspired by ranking principles. LabelCoRank leverages label co-occurrence relationships to refine initial label classifications through a dual-stage reranking process. The first stage uses initial classification results to form a preliminary ranking. In the second stage, a label co-occurrence matrix is utilized to rerank the preliminary results, enhancing the accuracy and relevance of the final classifications. By integrating the reranked label representations as additional text features, LabelCoRank effectively mitigates long tail issues in multi-labeltext classification. Experimental evaluations on popular datasets including MAG-CS, PubMed, and AAPD demonstrate the effectiveness and robustness of LabelCoRank.

多标签分类长尾问题标签共现重排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。