用大模型自动提取荷兰脑部MRI报告结构化信息,准确率超90%。
Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model
- 用LLaMA 3.1大模型零样本提取报告中的病灶评分与描述,支持多语言输入。
- 对萎缩和微出血等关键指标提取准确率达82%-96%,但数量统计误差较大。
- 少量示例提示可显著提升数量类信息提取效果,适合医学数据自动化研究者。
目的:从自由文本放射科报告中自动提取数据有助于大规模研究,但鲜有研究评估大语言模型在荷兰神经放射报告上的表现。方法:分析了来自三级记忆门诊的947份脑部MRI报告(2016-2021年),由主治神经放射科医生撰写。培训医学生标注了30个变量;100份报告双人标注以评估评分者间一致性。评估了开放权重大模型LLaMA 3.1在不同语言(荷兰语与英文翻译)及少样本提示下不同示例选择策略的表现。性能通过分类变量的平衡准确率、计数变量的准确率与平均绝对误差、自由文本的文本相似度进行评估,结果基于947份报告的10次随机划分计算。结果:LLaMA 3.1在视觉评分上表现出高零样本性能:左侧海马旁萎缩90%[77-100%],右侧96%[94-99%],全皮质萎缩87%[83-91%],Fazekas评分94%[93-96%]。微出血提及准确率达93%[92-95%],梗死提及为82%[80-84%]。病灶位置文本相似度达0.95[0.95-0.96]。数值类变量表现较差:微出血数量为80%[78-82%],梗死数量为66%[63-68%]。英文翻译结果相当。少样本提示提升了数值类变量表现,采用结构相似性选择示例后,微出血达92%[90-93%],梗死达81%[77-85%]。结论:LLaMA 3.1在提取荷兰神经放射报告方面具有强潜力。少样本提示能有效提升数值类信息提取,但位置相关变量仍存挑战。
原文摘要 · Abstract (English)
Objectives: Automatic data extraction from free-text radiology reports enables large-scale research, but few studies assessed the performance of large language models (LLMs) on Dutch neuroradiology reports. Methods: We analyzed 947 brain MRI reports from a tertiary memory clinic (2016-2021), authored by consultant neuroradiologists. Trained medical students annotated thirty variables; 100 reports were double-annotated to assess inter-rater reliability. We evaluated the performance of the open-weight LLM LLaMA 3.1 using different languages (Dutch vs. English translation) and few-shot prompting with different example selection strategies. Performance was evaluated using balanced accuracy for categorical variables, accuracy and mean absolute error for counts, and text similarity for free-text. Metrics were computed across 10 random splits of the 947 reports. Results: LLaMA 3.1 demonstrated high zero-shot performance for visual rating scores (mean [95%-CI]): Medial Temporal Atrophy: 90% [77-100%] on the left and 96% [94-99%] on the right, Global Cortical Atrophy: 87% [83-91%], and Fazekas: 94% [93-96%]. Microbleed mentions were detected with 93% accuracy [92-95%] and infarct mentions with 82% [80-84%]. Text similarity for lesion location reached 0.95 [0.95-0.96]. Performance was lower for numerical variables: 80% [78-82%] for the number of microbleeds and 66% [63-68%] for infarcts. English translation yielded comparable results. Few-shot prompting improved performance for numerical variables, achieving 92% [90-93%] for microbleeds and 81% [77-85%] for infarcts using structural similarity-based selection. Conclusion: LLaMA 3.1 shows strong potential for extracting data from Dutch neuroradiology reports. Few-shot prompting enhances performance for numerical variables, whereas challenges remain for location-specific variables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。