arXiv:2412.12744cs.CLcs.AI2024-12

跨领域对比文本分类方法,发现外域技术可带来新突破

Your Next State-of-the-Art Could Come from Another Domain: A Cross-Domain Analysis of Hierarchical Text Classification

  • 构建统一框架,横向对比不同领域的文本分类方法
  • 跨域应用技术,在评估中刷新多项指标纪录
  • 为多领域文本分类提供可复用的优化思路

层次化文本分类在自然语言处理中广泛存在且具有挑战性,例如为病历分配ICD编码、为专利打标IPC类别、为欧盟法律文本标注EUROVOC描述符等。尽管应用广泛,但对各领域最先进方法的系统性理解仍不足。本文首次提供全面的跨领域实证分析,提出统一框架将各类方法置于共同结构中进行比较。通过统一评估流程,我们验证了跨领域学习的必要性,并发现将其他领域的先进技术迁移应用,可在多个任务上取得新最优结果。

原文摘要 · Abstract (English)

Text classification with hierarchical labels is a prevalent and challenging task in natural language processing. Examples include assigning ICD codes to patient records, tagging patents into IPC classes, assigning EUROVOC descriptors to European legal texts, and more. Despite its widespread applications, a comprehensive understanding of state-of-the-art methods across different domains has been lacking. In this paper, we provide the first comprehensive cross-domain overview with empirical analysis of state-of-the-art methods. We propose a unified framework that positions each method within a common structure to facilitate research. Our empirical analysis yields key insights and guidelines, confirming the necessity of learning across different research areas to design effective methods. Notably, under our unified evaluation pipeline, we achieved new state-of-the-art results by applying techniques beyond their original domains.

文本分类跨领域层次化实证分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。