arXiv:2501.09214cs.CL2025-01AAAI被引 10

通过多源信息与双层级对比学习提升短文本分类效果

Boosting Short Text Classification with Multi-Source Information Exploration and Dual-Level Contrastive Learning

  • 融合统计、语言和事实三类信息缓解语义稀疏
  • 在多个数据集上超越主流模型,部分超过大语言模型
  • 设计分层结构显式建模任务间关联,提升学习效率

短文本分类因语义稀疏和标注样本不足而更具挑战性。本文提出新型模型MI-DELIGHT,首先探索统计、语言和事实三类多源信息以缓解语义稀疏问题;随后采用图学习方法对短文本进行建模,将其表示为图结构;进一步引入实例级与聚类级双层级对比学习辅助任务,有效挖掘大量无标签数据中的细粒度对比信息。以往模型仅并行执行主任务与辅助任务,未考虑任务间关联,因此本文设计分层架构,显式建模任务间相关性。在多个基准数据集上的大量实验表明,MI-DELIGHT显著优于现有先进模型,甚至在部分数据集上超越主流大语言模型。

原文摘要 · Abstract (English)

Short text classification, as a research subtopic in natural language processing, is more challenging due to its semantic sparsity and insufficient labeled samples in practical scenarios. We propose a novel model named MI-DELIGHT for short text classification in this work. Specifically, it first performs multi-source information (i.e., statistical information, linguistic information, and factual information) exploration to alleviate the sparsity issues. Then, the graph learning approach is adopted to learn the representation of short texts, which are presented in graph forms. Moreover, we introduce a dual-level (i.e., instance-level and cluster-level) contrastive learning auxiliary task to effectively capture different-grained contrastive information within massive unlabeled data. Meanwhile, previous models merely perform the main task and auxiliary tasks in parallel, without considering the relationship among tasks. Therefore, we introduce a hierarchical architecture to explicitly model the correlations between tasks. We conduct extensive experiments across various benchmark datasets, demonstrating that MI-DELIGHT significantly surpasses previous competitive models. It even outperforms popular large language models on several datasets.

短文本分类对比学习多源信息图学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。