用自研大模型从放射科报告中精准提取鉴别诊断,效果媲美GPT-4
Fine-Tuning In-House Large Language Models to Infer Differential Diagnosis from Radiology Reports
- 用GPT-4生成3.1万份标注数据,微调开源模型识别诊断
- 在1067份临床标注报告上达到92.1%的F1分数,接近GPT-4表现
- 适合医疗单位构建私有化模型,保障患者隐私与成本控制
放射科报告包含影像检查的关键发现和鉴别诊断信息,其提取对患者管理与治疗规划至关重要。然而,报告语言风格多样、格式不一,导致信息抽取困难。尽管专有大模型如GPT-4能有效获取临床信息,但高昂成本与受保护健康信息(PHI)隐私风险限制了实际应用。本研究提出一套定制化内部LLM开发流程,利用GPT-4生成31,056份标注报告后,微调开源模型以识别鉴别诊断。在由临床医生标注的1,067份报告上,该模型平均F1得分为92.1%,与GPT-4的90.8%相当。本方法为医疗机构提供了可替代昂贵专有模型的方案,实现性能相当、成本更低、隐私更安全的本地化诊断辅助。
原文摘要 · Abstract (English)
Radiology reports summarize key findings and differential diagnoses derived from medical imaging examinations. The extraction of differential diagnoses is crucial for downstream tasks, including patient management and treatment planning. However, the unstructured nature of these reports, characterized by diverse linguistic styles and inconsistent formatting, presents significant challenges. Although proprietary large language models (LLMs) such as GPT-4 can effectively retrieve clinical information, their use is limited in practice by high costs and concerns over the privacy of protected health information (PHI). This study introduces a pipeline for developing in-house LLMs tailored to identify differential diagnoses from radiology reports. We first utilize GPT-4 to create 31,056 labeled reports, then fine-tune open source LLM using this dataset. Evaluated on a set of 1,067 reports annotated by clinicians, the proposed model achieves an average F1 score of 92.1\%, which is on par with GPT-4 (90.8\%). Through this study, we provide a methodology for constructing in-house LLMs that: match the performance of GPT, reduce dependence on expensive proprietary models, and enhance the privacy and security of PHI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。