用微调大模型自动提取胰腺囊肿特征并分类,准确率媲美GPT-4o。
Leveraging Fine-Tuned Large Language Models for Interpretable Pancreatic Cystic Lesion Feature Extraction and Risk Categorization
- 用链式思维提示训练开源大模型,实现可解释的特征抽取。
- 特征提取准确率达97%,风险分类F1值达0.95,接近GPT-4o水平。
- 结果与放射科医生一致,适合大规模胰腺囊肿研究使用。
背景:手动从影像报告中提取胰腺囊肿(PCL)特征耗时费力,限制了推进PCL研究所需的大规模研究。目的:开发并评估能够自动从MRI/CT报告中提取PCL特征并依据指南分配风险等级的大语言模型(LLM)。方法:我们收集了2005–2024年间来自5,134名患者的6,000份腹部MRI/CT报告,描述了PCL。标签由GPT-4o通过链式思维(CoT)提示生成,用于提取PCL及主胰管特征。使用QLoRA在GPT-4o生成的CoT数据上微调了两个开源LLM。根据2017年ACR白皮书指南将特征映射为风险类别。在285份保留的人工标注报告上进行评估。对100例模型输出由三位放射科医生独立评审。特征提取采用精确匹配准确率评估,风险分类采用宏平均F1分数,放射科医生与模型一致性采用Fleiss' Kappa评估。结果:链式思维微调显著提升特征提取准确率:LLaMA从80%升至97%,DeepSeek从79%升至98%,与GPT-4o(97%)相当。风险分类F1得分也提升(LLaMA: 0.95;DeepSeek: 0.94),接近GPT-4o(0.97),无统计学差异。放射科医生间一致性高(Fleiss' Kappa = 0.888),加入DeepSeek-FT-CoT(Fleiss' Kappa = 0.893)或GPT-CoT(Fleiss' Kappa = 0.897)后仍无显著差异,表明两模型表现与放射科医生相当。
原文摘要 · Abstract (English)
Background: Manual extraction of pancreatic cystic lesion (PCL) features from radiology reports is labor-intensive, limiting large-scale studies needed to advance PCL research. Purpose: To develop and evaluate large language models (LLMs) that automatically extract PCL features from MRI/CT reports and assign risk categories based on guidelines. Materials and Methods: We curated a training dataset of 6,000 abdominal MRI/CT reports (2005-2024) from 5,134 patients that described PCLs. Labels were generated by GPT-4o using chain-of-thought (CoT) prompting to extract PCL and main pancreatic duct features. Two open-source LLMs were fine-tuned using QLoRA on GPT-4o-generated CoT data. Features were mapped to risk categories per institutional guideline based on the 2017 ACR White Paper. Evaluation was performed on 285 held-out human-annotated reports. Model outputs for 100 cases were independently reviewed by three radiologists. Feature extraction was evaluated using exact match accuracy, risk categorization with macro-averaged F1 score, and radiologist-model agreement with Fleiss' Kappa. Results: CoT fine-tuning improved feature extraction accuracy for LLaMA (80% to 97%) and DeepSeek (79% to 98%), matching GPT-4o (97%). Risk categorization F1 scores also improved (LLaMA: 0.95; DeepSeek: 0.94), closely matching GPT-4o (0.97), with no statistically significant differences. Radiologist inter-reader agreement was high (Fleiss' Kappa = 0.888) and showed no statistically significant difference with the addition of DeepSeek-FT-CoT (Fleiss' Kappa = 0.893) or GPT-CoT (Fleiss' Kappa = 0.897), indicating that both models achieved agreement levels on par with radiologists. Conclusion: Fine-tuned open-source LLMs with CoT supervision enable accurate, interpretable, and efficient phenotyping for large-scale PCL research, achieving performance comparable to GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。