arXiv:2504.19021cs.CLcs.AI2025-04被引 5

用扩展数据+投票策略,提升科学文本分类准确率

Advancing Scientific Text Classification: Fine-Tuned Models with Dataset Expansion and Hard-Voting

  • 用检索和模型预测扩充数据,生成1000条每类的新样本
  • 硬投票融合多模型预测,准确率显著提升
  • 专用模型在细分领域表现更优,适合学术自动化场景

高效文本分类对处理日益增长的学术出版物至关重要。本研究探索了BERT、SciBERT、BioBERT和BlueBERT等预训练语言模型(PLMs)在Web of Science(WoS-46985)数据集上的微调应用。为提升性能,通过在WoS数据库中执行七项定向查询,按WoS-46985主要类别各获取1000篇文献,构建新增数据集。利用PLMs对这些未标注数据进行标签预测,并采用硬投票策略融合结果以提高准确性和置信度。在扩展数据集上使用动态学习率与早停机制进行微调,显著提升了分类准确率,尤其在专业领域表现突出。领域专用模型如SciBERT和BioBERT持续优于通用模型BERT。结果表明,数据增强、基于推理的标签预测、硬投票及微调技术能有效构建鲁棒且可扩展的自动化学术文本分类系统。

原文摘要 · Abstract (English)

Efficient text classification is essential for handling the increasing volume of academic publications. This study explores the use of pre-trained language models (PLMs), including BERT, SciBERT, BioBERT, and BlueBERT, fine-tuned on the Web of Science (WoS-46985) dataset for scientific text classification. To enhance performance, we augment the dataset by executing seven targeted queries in the WoS database, retrieving 1,000 articles per category aligned with WoS-46985's main classes. PLMs predict labels for this unlabeled data, and a hard-voting strategy combines predictions for improved accuracy and confidence. Fine-tuning on the expanded dataset with dynamic learning rates and early stopping significantly boosts classification accuracy, especially in specialized domains. Domain-specific models like SciBERT and BioBERT consistently outperform general-purpose models such as BERT. These findings underscore the efficacy of dataset augmentation, inference-driven label prediction, hard-voting, and fine-tuning techniques in creating robust and scalable solutions for automated academic text classification.

文本分类科学文本数据增强硬投票

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。