用大模型理解医生开单文字,自动推荐精准胸部CT扫描方案。
Automated Chest CT Protocol Selection via Large Language Model Derived Text Embeddings from Imaging Request Text
- 用微调的大语言模型提取病历文本特征,捕捉临床描述的语义细节。
- 在18类协议上准确率达79%,与放射科医生水平相当(p=0.263)。
- 适合需要理解自由文本的医疗影像流程自动化场景。
目的:准确选择CT扫描方案对诊断质量与患者安全至关重要,但当前流程依赖人工,耗时且易出错。以往基于关键词或词袋的机器学习方法缺乏上下文理解能力,对罕见方案表现差。本文提出一种基于大语言模型(LLM)特征的决策支持系统,通过分析自由文本的临床指征,推荐胸部CT协议,以捕捉临床语义和表述差异,提升选择一致性与效率。方法:本研究获伦理委员会批准,回顾性分析某大型教学医院(2017–2024年)285,123份胸部CT申请,按80%/20%分为训练集(228,099例)与测试集(57,024例)。每份请求包含检查名称、临床指征、HIS备注及实际选择的协议。采用微调后的元公司大模型LLaMA-3.1-70B对临床文本进行嵌入,输入逻辑回归分类器预测18种协议标签(如PE、LDCT)。结果:该流程在18类协议上实现加权精确率0.84、加权F1分数0.81、总体准确率79%。在300例独立病例中,经专家共识验证,模型总体准确率为80%,放射科医生为83%,无显著差异(p=0.263)。各类别表现接近,部分困难类别中模型优于医生;熵分析显示协议使用更均衡,表明变异性降低。结论:基于大语言模型的推荐系统能利用大规模自然文本语料中的通用知识,从自由文本请求中准确匹配胸部CT协议,可作为需语言理解的协议推荐工具的可行基础。
原文摘要 · Abstract (English)
Purpose: Accurate CT protocol selection is critical for diagnostic quality and patient safety, yet the current process is manual, time-consuming, and prone to inconsistencies. Prior Machine Learning methods using keywords or bag-of-words lack contextual understanding and perform poorly on rare protocols. We propose a decision support system using large language model (LLM) features to recommend protocols from free-text clinical indications, capturing clinical nuance and phrasing variation for more consistent, efficient selection. Methods: In this REB-approved retrospective study, 285,123 chest CT imaging requests from a large academic medical center (2017-2024) were split into training (228,099, 80%) and held-out test (57,024, 20%) sets. Each request included procedure names, clinical indication, HIS comments, and the selected protocol. Clinical text was embedded using a fine-tuned LLM, Meta's LLaMA-3.1-70B; these features input a logistic regression classifier predicting 18 protocol labels (e.g., PE, LDCT). Results: The pipeline achieved a weighted precision of 0.84, weighted F1-score of 0.81, and overall accuracy of 79% across 18 CT protocols. On 300 independent cases with expert consensus, the LLM reached an overall accuracy of 80% versus 83% for radiologists, with no significant difference (p = 0.263). Performance was comparable across most classes, with the LLM exceeding radiologists for some challenging categories, and entropy analyses indicated more balanced protocol use, suggesting reduced variability. Conclusion: An LLM-based recommendation system can leverage general knowledge from a large natural-text corpus to accurately assign chest CT protocols from free-text imaging requests, and may serve as a viable foundation for protocol recommendation tools where inputs require language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。