九款开源大模型在心磁报告诊断提取中表现优异,最高F1达0.98。
Comparative analysis of privacy-preserving open-source LLMs regarding extraction of diagnostic information from clinical CMR imaging reports
- 对比九款本地部署的开源大模型,评估其从心磁报告中提取诊断信息的能力。
- 谷歌Gemma2模型F1得分0.98,领先其他模型,均超0.93,部分超越专业医生。
- 结果支持在临床中使用开源模型实现快速、隐私保护的自动化诊断分类。
目的:研究隐私保护、本地部署的开源大型语言模型(LLMs)在从自由文本心血管磁共振(CMR)报告中提取诊断信息的应用。方法:在109份临床CMR报告上评估九款开源LLMs识别诊断并分类患者至不同心脏诊断类别能力,使用准确率、精确率、召回率和F1分数等标准分类指标量化性能,并通过混淆矩阵分析各模型的误分类模式。结果:多数开源LLMs在报告分类中表现卓越。谷歌Gemma2模型平均F1得分为0.98,居首;Qwen2.5:32B与DeepseekR1-32B分别达0.96和0.95;其余模型平均分均高于0.93,仅Mistral与DeepseekR1-7B例外。前四名模型在所有评估指标上均优于资深心脏病专家(F1=0.94)。结论:研究证实开源、隐私保护型LLMs在临床环境中用于影像报告自动分析的可行性,可实现高效、精准、低资源消耗的诊断分类。
原文摘要 · Abstract (English)
Purpose: We investigated the utilization of privacy-preserving, locally-deployed, open-source Large Language Models (LLMs) to extract diagnostic information from free-text cardiovascular magnetic resonance (CMR) reports. Materials and Methods: We evaluated nine open-source LLMs on their ability to identify diagnoses and classify patients into various cardiac diagnostic categories based on descriptive findings in 109 clinical CMR reports. Performance was quantified using standard classification metrics including accuracy, precision, recall, and F1 score. We also employed confusion matrices to examine patterns of misclassification across models. Results: Most open-source LLMs demonstrated exceptional performance in classifying reports into different diagnostic categories. Google's Gemma2 model achieved the highest average F1 score of 0.98, followed by Qwen2.5:32B and DeepseekR1-32B with F1 scores of 0.96 and 0.95, respectively. All other evaluated models attained average scores above 0.93, with Mistral and DeepseekR1-7B being the only exceptions. The top four LLMs outperformed our board-certified cardiologist (F1 score of 0.94) across all evaluation metrics in analyzing CMR reports. Conclusion: Our findings demonstrate the feasibility of implementing open-source, privacy-preserving LLMs in clinical settings for automated analysis of imaging reports, enabling accurate, fast and resource-efficient diagnostic categorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。