arXiv:2510.05664cs.AI2025-10被引 2

用GPT-4o从放射报告中提取带不确定性的诊断标签,训练出高性能多标签分类模型。

Large Language Model-Based Uncertainty-Adjusted Label Extraction for Artificial Intelligence Model Development in Upper Extremity Radiography

  • 让GPT-4o自动识别报告中的影像发现并标注为存在、不存在或不确定。
  • 标签提取准确率达98.6%,模型在多个数据集上AUC均超0.79。
  • 即使考虑不确定性,模型性能仍稳定,适合医疗影像自动化标注。

目的:评估GPT-4o从自由文本放射科报告中提取诊断标签(含不确定性)的能力,并测试这些标签对上肢骨骼放射影像多标签分类的影响。方法:本回顾性研究包含锁骨(n=1,170)、肘部(n=3,755)和拇指(n=1,978)的放射影像系列。匿名化后,GPT-4o通过结构化模板标记影像发现为“存在”(真)、“不存在”(假)或“不确定”。为评估标签不确定性的影响,训练和验证集中的“不确定”标签被自动重分配为“存在”(包容)或“不存在”(排他)。使用ResNet50对标签-图像对进行多标签分类。在内部(锁骨:n=233,肘部:n=745,拇指:n=393)和外部测试集(每类n=300)上人工验证标签提取准确性。性能通过宏平均受试者工作特征曲线下面积(AUC)、精确率-召回率曲线、敏感性、特异性及准确率评估。采用DeLong检验比较AUC差异。结果:测试集中自动提取的标签正确率为98.6%(60,618/61,488)。在不同解剖区域,基于标签的模型训练取得具有竞争力的性能,包容型模型的宏平均AUC如肘部达0.80(范围0.62–0.87),排他型模型为0.80(范围0.61–0.88)。模型在外部数据集上表现良好(肘部[包容]:AUC=0.79,范围0.61–0.87;肘部[排他]:AUC=0.79,范围0.63–0.89)。不同标注策略或数据集间无显著差异(p≥0.15)。结论:GPT-4o可从放射报告中提取标签以训练高性能多标签分类模型,且检测到的不确定性未影响模型性能。

原文摘要 · Abstract (English)

Objectives: To evaluate GPT-4o's ability to extract diagnostic labels (with uncertainty) from free-text radiology reports and to test how these labels affect multi-label image classification of musculoskeletal radiographs. Methods: This retrospective study included radiography series of the clavicle (n=1,170), elbow (n=3,755), and thumb (n=1,978). After anonymization, GPT-4o filled out structured templates by indicating imaging findings as present ("true"), absent ("false"), or "uncertain." To assess the impact of label uncertainty, "uncertain" labels of the training and validation sets were automatically reassigned to "true" (inclusive) or "false" (exclusive). Label-image-pairs were used for multi-label classification using ResNet50. Label extraction accuracy was manually verified on internal (clavicle: n=233, elbow: n=745, thumb: n=393) and external test sets (n=300 for each). Performance was assessed using macro-averaged receiver operating characteristic (ROC) area under the curve (AUC), precision recall curves, sensitivity, specificity, and accuracy. AUCs were compared with the DeLong test. Results: Automatic extraction was correct in 98.6% (60,618 of 61,488) of labels in the test sets. Across anatomic regions, label-based model training yielded competitive performance measured by macro-averaged AUC values for inclusive (e.g., elbow: AUC=0.80 [range, 0.62-0.87]) and exclusive models (elbow: AUC=0.80 [range, 0.61-0.88]). Models generalized well on external datasets (elbow [inclusive]: AUC=0.79 [range, 0.61-0.87]; elbow [exclusive]: AUC=0.79 [range, 0.63-0.89]). No significant differences were observed across labeling strategies or datasets (p>=0.15). Conclusion: GPT-4o extracted labels from radiologic reports to train competitive multi-label classification models with high accuracy. Detected uncertainty in the radiologic reports did not influence the performance of these models.

大模型医学影像多标签分类自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。