arXiv:2504.21191cs.CLcs.AI2025-04被引 11

小模型微调后比大模型零样本更适配医疗文本分类。

Small or Large? Zero-Shot or Finetuned? Guiding Language Model Choice for Specialized Applications in Healthcare

  • 小模型经微调显著优于零样本使用,尤其在复杂任务上
  • 领域相关小模型微调后表现超越通用大模型零样本
  • 小模型在数据稀缺时仍具优势,适合资源受限场景

本研究通过分析不列颠哥伦比亚癌症登记库(BCCR)的电子病理报告,评估了三种不同难度和数据规模的分类任务。比较了多种小语言模型(SLMs)与一个大语言模型(LLM)的表现,其中小模型采用零样本和微调两种方式,大模型仅零样本。结果显示:微调使小模型在所有任务中性能大幅提升;零样本大模型虽优于零样本小模型,但始终不及微调后的小模型。经过微调后,领域邻近小模型在困难任务上表现更佳。进一步领域预训练在简单任务中带来小幅提升,在复杂且数据稀缺任务中实现显著改进。研究表明,小模型通过微调可超越大模型零样本表现,领域相关或特定预训练能进一步提升效果。尽管大模型具备强大零样本能力,但在特定任务中仍不及合理微调的小模型。在大模型时代,小模型依然有效,提供更优的性能与资源平衡。

原文摘要 · Abstract (English)

This study aims to guide language model selection by investigating: 1) the necessity of finetuning versus zero-shot usage, 2) the benefits of domain-adjacent versus generic pretrained models, 3) the value of further domain-specific pretraining, and 4) the continued relevance of Small Language Models (SLMs) compared to Large Language Models (LLMs) for specific tasks. Using electronic pathology reports from the British Columbia Cancer Registry (BCCR), three classification scenarios with varying difficulty and data size are evaluated. Models include various SLMs and an LLM. SLMs are evaluated both zero-shot and finetuned; the LLM is evaluated zero-shot only. Finetuning significantly improved SLM performance across all scenarios compared to their zero-shot results. The zero-shot LLM outperformed zero-shot SLMs but was consistently outperformed by finetuned SLMs. Domain-adjacent SLMs generally performed better than the generic SLM after finetuning, especially on harder tasks. Further domain-specific pretraining yielded modest gains on easier tasks but significant improvements on the complex, data-scarce task. The results highlight the critical role of finetuning for SLMs in specialized domains, enabling them to surpass zero-shot LLM performance on targeted classification tasks. Pretraining on domain-adjacent or domain-specific data provides further advantages, particularly for complex problems or limited finetuning data. While LLMs offer strong zero-shot capabilities, their performance on these specific tasks did not match that of appropriately finetuned SLMs. In the era of LLMs, SLMs remain relevant and effective, offering a potentially superior performance-resource trade-off compared to LLMs.

小模型医疗AI微调领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。