arXiv:2412.00098cs.CL2024-12被引 23

用科学领域模型提升论文分类准确率

Fine-Tuning Large Language Models for Scientific Text Classification: A Comparative Study

  • 在三个科学文本数据集上微调四种大模型
  • 领域专用模型SciBERT表现优于通用模型
  • 适合需要高精度学术文本分类的研究者

在线文本内容的指数级增长促使自动化文本分类方法的发展。基于Transformer架构的大语言模型(LLM)在自然语言处理任务中表现出色,但通用模型在科学文本等专业领域面临术语复杂、数据不平衡等问题。本研究在来自WoS-46985数据集的三个数据集上,对四种先进模型BERT、SciBERT、BioBERT和BlueBERT进行微调,评估其在科学文本分类中的表现。实验表明,领域专用模型(尤其是SciBERT)在基于摘要和关键词的分类任务中均显著优于通用模型。此外,与文献中报道的深度学习模型结果对比,进一步凸显了大语言模型在特定领域的优势。研究强调了针对领域特性优化大模型的重要性,以提升其在专业文本分类中的有效性。

原文摘要 · Abstract (English)

The exponential growth of online textual content across diverse domains has necessitated advanced methods for automated text classification. Large Language Models (LLMs) based on transformer architectures have shown significant success in this area, particularly in natural language processing (NLP) tasks. However, general-purpose LLMs often struggle with domain-specific content, such as scientific texts, due to unique challenges like specialized vocabulary and imbalanced data. In this study, we fine-tune four state-of-the-art LLMs BERT, SciBERT, BioBERT, and BlueBERT on three datasets derived from the WoS-46985 dataset to evaluate their performance in scientific text classification. Our experiments reveal that domain-specific models, particularly SciBERT, consistently outperform general-purpose models in both abstract-based and keyword-based classification tasks. Additionally, we compare our achieved results with those reported in the literature for deep learning models, further highlighting the advantages of LLMs, especially when utilized in specific domains. The findings emphasize the importance of domain-specific adaptations for LLMs to enhance their effectiveness in specialized text classification tasks.

大模型微调文本分类科学文本领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。