arXiv:2411.19119cs.IRcs.LG2024-11被引 1

新推出三个科研论文层级分类数据集,提升基准测试可靠性。

Introducing Three New Benchmark Datasets for Hierarchical Text Classification

  • 基于期刊与引用分类法融合构建更可靠的层级标签体系。
  • 新数据集内同类文档语义相似度显著高于原有基准。
  • 为科学文献分类研究提供可复现的性能基线,适合算法评估者使用。

层级文本分类(HTC)旨在将文本文档归入具有结构化层级关系的类别中。现有主流方法多依赖三个经典基准数据集:Web of Science(WOS)、Reuters Corpus Volume 1 Version 2(RCV1-V2)和New York Times(NYT)。但除RCV1-V2外,其余数据集缺乏详细标注方法说明。本文在科研出版领域引入三个新HTC数据集,其内容来自WOS数据库的论文标题与摘要。首先构建两个基于现有期刊和引用分类体系的基准数据集;因二者各有缺陷,提出融合策略以提升分类可靠性和鲁棒性。通过聚类分析验证,融合方案生成的数据集内同类文档语义更相近。最后,对四种先进HTC模型在三组新数据集上的表现进行评估,为未来科学文献分类的机器学习方法研究提供基准参考。

原文摘要 · Abstract (English)

Hierarchical Text Classification (HTC) is a natural language processing task with the objective to classify text documents into a set of classes from a structured class hierarchy. Many HTC approaches have been proposed which attempt to leverage the class hierarchy information in various ways to improve classification performance. Machine learning-based classification approaches require large amounts of training data and are most-commonly compared through three established benchmark datasets, which include the Web Of Science (WOS), Reuters Corpus Volume 1 Version 2 (RCV1-V2) and New York Times (NYT) datasets. However, apart from the RCV1-V2 dataset which is well-documented, these datasets are not accompanied with detailed description methodologies. In this paper, we introduce three new HTC benchmark datasets in the domain of research publications which comprise the titles and abstracts of papers from the Web of Science publication database. We first create two baseline datasets which use existing journal-and citation-based classification schemas. Due to the respective shortcomings of these two existing schemas, we propose an approach which combines their classifications to improve the reliability and robustness of the dataset. We evaluate the three created datasets with a clustering-based analysis and show that our proposed approach results in a higher quality dataset where documents that belong to the same class are semantically more similar compared to the other datasets. Finally, we provide the classification performance of four state-of-the-art HTC approaches on these three new datasets to provide baselines for future studies on machine learning-based techniques for scientific publication classification.

层级分类科研文献数据集基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。