arXiv:2503.15169cs.CLcs.AI2025-03

对比5个开源大模型在6项医疗文本分类任务中的表现,发现模型大小不决定性能。

Benchmarking Open-Source Large Language Models on Healthcare Text Classification Tasks

  • 在社交媒体与临床数据上测试5个开源大模型的分类能力
  • DeepSeekV3在4项任务中表现最优,但小模型GEMMA召回率很高
  • 模型性能差异大,选型需看任务类型而非参数规模

大语言模型在医疗信息提取中的应用日益受到关注。本研究评估了五个开源大模型(GEMMA-3-27B-IT、LLAMA3-70B、LLAMA4-109B、DEEPSEEK-R1-DISTILL-LLAMA-70B、DEEPSEEK-V3-0324-UD-Q2_K_XL)在六项医疗相关分类任务上的表现,涵盖社交媒体数据(乳腺癌、用药变化、妊娠不良结局、潜在新冠病例)和临床数据(污名化标签、用药讨论)。报告所有模型-任务组合的精确率、召回率和F1分数及其95%置信区间。结果显示模型间性能差异显著,DeepSeekV3总体表现最佳,在四项任务中取得最高F1值。模型在社交媒体任务上整体优于临床数据任务,表明存在领域特异性挑战。尽管参数量较小,GEMMA-3-27B-IT表现出极高的召回率;而LLAMA4-109B的表现远逊于其前代模型LLAMA3-70B,说明参数规模并非性能保障。各模型呈现不同精度-召回权衡,部分侧重敏感性,部分侧重特异性。结果强调在医疗应用中应基于具体数据域和精度-召回需求进行模型选择,而非仅依赖模型规模。随着医疗领域越来越多采用AI驱动的文本分类工具,本次全面基准测试为模型选型与实施提供了重要参考,也凸显了持续评估与领域适配的必要性。

原文摘要 · Abstract (English)

The application of large language models (LLMs) to healthcare information extraction has emerged as a promising approach. This study evaluates the classification performance of five open-source LLMs: GEMMA-3-27B-IT, LLAMA3-70B, LLAMA4-109B, DEEPSEEK-R1-DISTILL-LLAMA-70B, and DEEPSEEK-V3-0324-UD-Q2_K_XL, across six healthcare-related classification tasks involving both social media data (breast cancer, changes in medication regimen, adverse pregnancy outcomes, potential COVID-19 cases) and clinical data (stigma labeling, medication change discussion). We report precision, recall, and F1 scores with 95% confidence intervals for all model-task combinations. Our findings reveal significant performance variability between LLMs, with DeepSeekV3 emerging as the strongest overall performer, achieving the highest F1 scores in four tasks. Notably, models generally performed better on social media tasks compared to clinical data tasks, suggesting potential domain-specific challenges. GEMMA-3-27B-IT demonstrated exceptionally high recall despite its smaller parameter count, while LLAMA4-109B showed surprisingly underwhelming performance compared to its predecessor LLAMA3-70B, indicating that larger parameter counts do not guarantee improved classification results. We observed distinct precision-recall trade-offs across models, with some favoring sensitivity over specificity and vice versa. These findings highlight the importance of task-specific model selection for healthcare applications, considering the particular data domain and precision-recall requirements rather than model size alone. As healthcare increasingly integrates AI-driven text classification tools, this comprehensive benchmarking provides valuable guidance for model selection and implementation while underscoring the need for continued evaluation and domain adaptation of LLMs in healthcare contexts.

大模型评测医疗文本分类任务模型选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。