arXiv:2506.03259cs.CL2025-06被引 2

用开源大模型零样本标注胸部腹部骨盆CT报告,效果优于传统方法。

Zero-Shot Multi-Disease Labeling of Chest, Abdomen, and Pelvis CT Reports Using Open-Weight Large Language Models: The Effect of Labeling Conventions

  • 用五种轻量级开源大模型零样本提示,自动标注多部位CT报告
  • 最大F1达0.84(集成模型),远超规则算法和微调的RadBERT
  • 标注习惯差异导致模型与医生分歧,建议统一标准以提升可比性

目的:比较五种轻量级开源大语言模型(LLMs)与基于规则的算法(RBA)及微调的RadBERT在零样本条件下对胸部-腹部-骨盆(CAP)CT报告进行15类疾病标签标注的效果,并研究标注规范对性能的影响。方法:回顾性分析2012至2017年间来自29,540名患者的40,833份CAP CT报告(年龄与性别信息缺失)。五种LLMs通过零样本提示生成标签,与RBA和微调后的RadBERT对比。在12,197份保留报告上计算模型间一致性(Cohen's kappa),并在1,789份放射科医师标注数据、简化版临床行动性忽略数据及CT-RATE数据集上计算宏平均F1分数。采用非重叠自助法构建95%置信区间判断显著差异。结果:MedGemma 27B与MedGemma-1.5 4B表现出最高中位一致率(κ = 0.90)。在人工标注下,Gemma-3 27B取得最高宏平均F1(0.82 [95% CI: 0.80, 0.83]),显著高于RadBERT(0.66)和RBA(0.64);多数投票集成模型达0.84。主观类标签表现最差,肾部病变(0.44)和肺不张(0.67)得分最低。重新标注后,所有模型在肾部病变上F1提升,但仅大模型在肺不张上提升,而RBA和RadBERT反而下降。所有模型在CT-RATE数据集上的F1均高于人工标注,反映其更注重字面表达。结论:轻量级开源大模型在零样本提示下显著优于规则算法和微调的BERT模型进行CAP CT报告标注。模型与医生的分歧主要源于不同的标注标准。

原文摘要 · Abstract (English)

Purpose: To compare five lightweight open-weight large language models (LLMs) with a rule-based algorithm (RBA) and fine-tuned RadBERT for zero-shot labeling of chest-abdomen-pelvis (CAP) CT reports, and to examine how labeling conventions affect measured performance. Materials and Methods: In this retrospective study, 40,833 CAP CT reports from 29,540 patients examined between 2012 and 2017 were analyzed; age and sex were unavailable. Five LLMs were prompted zero-shot to assign 15 labels across three organ systems and compared with an RBA and fine-tuned RadBERT. Inter-model agreement was assessed with Cohen kappa ($κ$) on 12,197 held-out reports. Macro-averaged F1 was computed against 1,789 radiologist-supervised annotations, the same annotations simplified to disregard clinical actionability, and the CT-RATE dataset. Nonoverlapping bootstrapped 95% CIs indicated relevant differences. Results: MedGemma 27B and MedGemma-1.5 4B showed the highest median agreement ($κ$ = 0.90). Against manual annotations, Gemma-3 27B achieved the highest macro-averaged F1 (0.82 [95% CI: 0.80, 0.83]) versus 0.66 for RadBERT and 0.64 for the RBA; a majority-vote ensemble scored 0.84. Scores were lowest averaged across models for subjective classes, kidney lesion (0.44) and atelectasis (0.67). Relabeling raised F1 for all models on kidney lesion, but for atelectasis only for the LLMs; the RBA and RadBERT declined. F1 against CT-RATE exceeded that against manual annotations for all models, reflecting its more literal convention. Conclusion: Lightweight open-weight LLMs outperformed rule-based and fine-tuned BERT labeling of CAP CT reports with zero-shot prompting. Models and annotators disagreed largely because they applied different labeling criteria.

医学影像大模型零样本自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。