用结构化标注和解剖定位提升大模型,精准识别需随访的意外发现。
Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches
- 引入病变标记与解剖提示,让大模型更准理解报告上下文。
- 最佳模型宏F1达0.79,超过所有监督模型(最高0.70),接近专家一致率0.76。
- 适合医学影像自动化筛查,尤其需要可解释性的临床场景。
目的:评估大语言模型(LLMs)在细粒度、病灶级检测需随访的意外发现方面是否优于传统监督模型,解决现有文档级分类系统的局限性。方法:使用包含1,623个验证病灶的400份标注放射科报告数据集,对比三种基于Transformer的监督编码器(BioClinicalModernBERT、ModernBERT、Clinical Longformer)与四种生成式LLM配置(Llama 3.1-8B、GPT-4o、GPT-OSS-20b)。提出新型推理策略,采用病灶标记输入与解剖感知提示以增强模型推理。性能通过类别特异性F1分数评估。结果:解剖信息增强的GPT-OSS-20b模型表现最优,获得0.79的意外发现阳性宏F1,超越所有监督基线(最高0.70),并接近0.76的标注者间一致性。显式解剖定位在所有GPT模型上带来统计显著的性能提升(p < 0.05),多数投票集成使宏F1进一步提升至0.90。错误分析显示,解剖感知的LLM在区分可行动病灶与良性病变方面具有更强的上下文推理能力。结论:经结构化病灶标记与解剖上下文增强的大语言模型,显著优于传统监督编码器,并达到与人类专家相当的性能。该方法为放射科工作流中的意外发现自动监测提供了可靠且可解释的路径。
原文摘要 · Abstract (English)
Objective: To evaluate large language models (LLMs) against supervised baselines for fine-grained, lesion-level detection of incidentalomas requiring follow-up, addressing the limitations of current document-level classification systems. Methods: We utilized a dataset of 400 annotated radiology reports containing 1,623 verified lesion findings. We compared three supervised transformer-based encoders (BioClinicalModernBERT, ModernBERT, Clinical Longformer) against four generative LLM configurations (Llama 3.1-8B, GPT-4o, GPT-OSS-20b). We introduced a novel inference strategy using lesion-tagged inputs and anatomy-aware prompting to ground model reasoning. Performance was evaluated using class-specific F1-scores. Results: The anatomy-informed GPT-OSS-20b model achieved the highest performance, yielding an incidentaloma-positive macro-F1 of 0.79. This surpassed all supervised baselines (maximum macro-F1: 0.70) and closely matched the inter-annotator agreement of 0.76. Explicit anatomical grounding yielded statistically significant performance gains across GPT-based models (p < 0.05), while a majority-vote ensemble of the top systems further improved the macro-F1 to 0.90. Error analysis revealed that anatomy-aware LLMs demonstrated superior contextual reasoning in distinguishing actionable findings from benign lesions. Conclusion: Generative LLMs, when enhanced with structured lesion tagging and anatomical context, significantly outperform traditional supervised encoders and achieve performance comparable to human experts. This approach offers a reliable, interpretable pathway for automated incidental finding surveillance in radiology workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。