系统梳理LLM在医疗文本分类中的应用与挑战
Large Language Models for Healthcare Text Classification: A Systematic Review
- 基于PRISMA指南整合65篇文献,覆盖多种医疗文本分类任务
- 发现隐私保护与术语复杂性是核心难点,模型性能受数据不平衡影响
- 适合关注医疗AI落地的研究者与临床决策支持系统开发者
大型语言模型(LLMs)已深刻改变自然语言处理领域的范式。在医疗领域,准确高效的文本分类对临床笔记分析、诊断编码等任务至关重要,而LLMs展现出巨大潜力。然而,文本分类长期面临标注成本高、数据不平衡及可扩展性差等挑战,医疗场景下还需应对患者隐私保护和医学术语复杂性的额外难题。大量研究尝试利用LLMs实现自动化医疗文本分类,并与传统基于嵌入的机器学习方法对比。现有综述或未聚焦文本分类,或不专精于医疗领域。本研究依据PRISMA指南,系统检索2018至2024年期间的Google Scholar、Scopus、PubMed、Science Direct等数据库,共纳入65篇相关研究。按分类类型(如二分类、多标签分类)、应用场景(如临床决策支持、公共卫生与舆论分析)、方法学、医疗文本类型及评估指标进行分类分析。结果揭示了当前文献中的关键空白,并提出未来可探索的研究方向。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have fundamentally transformed approaches to Natural Language Processing (NLP) tasks across diverse domains. In healthcare, accurate and cost-efficient text classification is crucial, whether for clinical notes analysis, diagnosis coding, or any other task, and LLMs present promising potential. Text classification has always faced multiple challenges, including manual annotation for training, handling imbalanced data, and developing scalable approaches. With healthcare, additional challenges are added, particularly the critical need to preserve patients' data privacy and the complexity of the medical terminology. Numerous studies have been conducted to leverage LLMs for automated healthcare text classification and contrast the results with existing machine learning-based methods where embedding, annotation, and training are traditionally required. Existing systematic reviews about LLMs either do not specialize in text classification or do not focus on the healthcare domain. This research synthesizes and critically evaluates the current evidence found in the literature regarding the use of LLMs for text classification in a healthcare setting. Major databases (e.g., Google Scholar, Scopus, PubMed, Science Direct) and other resources were queried, which focused on the papers published between 2018 and 2024 within the framework of PRISMA guidelines, which resulted in 65 eligible research articles. These were categorized by text classification type (e.g., binary classification, multi-label classification), application (e.g., clinical decision support, public health and opinion analysis), methodology, type of healthcare text, and metrics used for evaluation and validation. This review reveals the existing gaps in the literature and suggests future research lines that can be investigated and explored.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。