arXiv:2510.17892cs.CL2025-10综述被引 21

系统梳理预训练模型在垂直领域文本分类中的应用与挑战。

Advances in Pre-trained Language Models for Domain-Specific Text Classification: A Systematic Review

  • 基于PRISMA标准筛选41篇2018-2024年文献,系统分析方法演进
  • 对比BERT、SciBERT、BioBERT在生物医学句子分类中表现
  • 提出技术分类体系,揭示大模型在专业领域应用的瓶颈

科学文献和在线信息的指数级增长亟需高效的知识提取方法。自然语言处理在文本分类任务中发挥关键作用。尽管大语言模型在通用NLP任务中取得显著进展,但在领域特定场景下,因专业词汇、独特语法结构及数据分布不均,准确率常受影响。本系统综述(SLR)聚焦预训练语言模型(PLMs)在领域文本分类中的应用,依据PRISMA声明,系统检索并分析2018年至2024年1月间发表的41篇论文。研究涵盖文本分类技术演进,区分传统与现代方法,重点分析基于Transformer的模型及其在领域适配中的挑战。我们按不同PLM分类现有研究,并构建技术分类体系。通过对比实验验证Bert、SciBERT与BioBERT在生物医学句子分类中的性能。进一步开展跨领域大模型分类性能比较。最后,总结近期进展,提出未来方向与局限性。

原文摘要 · Abstract (English)

The exponential increase in scientific literature and online information necessitates efficient methods for extracting knowledge from textual data. Natural language processing (NLP) plays a crucial role in addressing this challenge, particularly in text classification tasks. While large language models (LLMs) have achieved remarkable success in NLP, their accuracy can suffer in domain-specific contexts due to specialized vocabulary, unique grammatical structures, and imbalanced data distributions. In this systematic literature review (SLR), we investigate the utilization of pre-trained language models (PLMs) for domain-specific text classification. We systematically review 41 articles published between 2018 and January 2024, adhering to the PRISMA statement (preferred reporting items for systematic reviews and meta-analyses). This review methodology involved rigorous inclusion criteria and a multi-step selection process employing AI-powered tools. We delve into the evolution of text classification techniques and differentiate between traditional and modern approaches. We emphasize transformer-based models and explore the challenges and considerations associated with using LLMs for domain-specific text classification. Furthermore, we categorize existing research based on various PLMs and propose a taxonomy of techniques used in the field. To validate our findings, we conducted a comparative experiment involving BERT, SciBERT, and BioBERT in biomedical sentence classification. Finally, we present a comparative study on the performance of LLMs in text classification tasks across different domains. In addition, we examine recent advancements in PLMs for domain-specific text classification and offer insights into future directions and limitations in this rapidly evolving domain.

文本分类预训练模型领域适配系统综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。