arXiv:2409.16563cs.AI2024-09被引 9

用合成标签微调轻量级大模型,提升放射科报告疾病检测能力

Enhancing disease detection in radiology reports through fine-tuning lightweight LLM on weak labels

  • 用合成标签微调 Llama 3.1-8B 模型,解决医疗数据标注难问题
  • 高质量合成标签下,疾病检测微 F1 达 0.91,优于原始教师标签
  • 适合缺乏标注数据的医疗 AI 研究者参考,尤其关注模型轻量化

尽管大语言模型(LLMs)在医学领域取得显著进展,但模型规模限制和缺乏特定队列的标注数据集仍是实际应用的障碍。本文研究通过合成标签对轻量级 LLM(如 Llama 3.1-8B)进行微调的潜力。通过联合训练两个任务的指令数据集,当任务特定合成标签质量较高(如由 GPT4-o 生成)时,微调后的 Llama 3.1-8B 在开放式疾病检测任务上达到微 F1 分数 0.91。而当合成标签质量较低(如来自 MIMIC-CXR 数据集)时,经校准后模型性能仍优于其噪声教师标签(微 F1 0.67 vs 0.63),表明该模型具备强大内在能力。结果证明,使用合成标签微调 LLM 具有前景,为医学领域 LLM 专业化提供新方向。

原文摘要 · Abstract (English)

Despite significant progress in applying large language models (LLMs) to the medical domain, several limitations still prevent them from practical applications. Among these are the constraints on model size and the lack of cohort-specific labeled datasets. In this work, we investigated the potential of improving a lightweight LLM, such as Llama 3.1-8B, through fine-tuning with datasets using synthetic labels. Two tasks are jointly trained by combining their respective instruction datasets. When the quality of the task-specific synthetic labels is relatively high (e.g., generated by GPT4- o), Llama 3.1-8B achieves satisfactory performance on the open-ended disease detection task, with a micro F1 score of 0.91. Conversely, when the quality of the task-relevant synthetic labels is relatively low (e.g., from the MIMIC-CXR dataset), fine-tuned Llama 3.1-8B is able to surpass its noisy teacher labels (micro F1 score of 0.67 v.s. 0.63) when calibrated against curated labels, indicating the strong inherent underlying capability of the model. These findings demonstrate the potential of fine-tuning LLMs with synthetic labels, offering a promising direction for future research on LLM specialization in the medical domain.

医疗AI轻量模型合成标签疾病检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。