arXiv:2509.15083cs.CV2025-09

评估AI肺分割模型在重症肺病患者中的表现,发现严重病例下精度显著下降。

Transplant-Ready? Evaluating AI Lung Segmentation Models in Candidates with Severe Lung Disease

  • 对比三种深度学习模型在不同病情和病理类型下的分割效果
  • 重度病例中体积相似度明显下降,所有模型性能均减弱
  • 适用于临床术前规划的AI模型需针对重症病例专门优化

本研究评估了公开可用的深度学习肺分割模型在符合移植条件患者中的表现,以确定其在不同疾病严重程度、病理类别及肺侧别下的性能,并识别影响肺移植术前规划应用的局限性。回顾性研究纳入2017至2019年间杜克大学健康系统32名接受胸部CT扫描的患者(共3,645个2D轴向切片),入选标准为存在两种及以上不同程度肺部病变。采用三种已开发的深度学习模型(Unet-R231、TotalSegmentator、MedSAM)进行肺分割,通过定量指标(体积相似度、Dice相似系数、豪斯多夫距离)和定性评估(四点临床可接受性评分)进行性能分析。Unet-R231在整体表现、不同严重程度及病理类别中均优于TotalSegmentator和MedSAM(p<0.05)。所有模型在从轻度到中重度病例间均出现显著性能下降,尤其在体积相似度方面(p<0.05),但肺侧别与病理类型间无显著差异。Unet-R231在所评估模型中提供最准确的自动化肺分割,TotalSegmentator紧随其后,但两者在中重度病例中表现显著下降,强调在严重病理情境下需对模型进行针对性微调。

原文摘要 · Abstract (English)

This study evaluates publicly available deep-learning based lung segmentation models in transplant-eligible patients to determine their performance across disease severity levels, pathology categories, and lung sides, and to identify limitations impacting their use in preoperative planning in lung transplantation. This retrospective study included 32 patients who underwent chest CT scans at Duke University Health System between 2017 and 2019 (total of 3,645 2D axial slices). Patients with standard axial CT scans were selected based on the presence of two or more lung pathologies of varying severity. Lung segmentation was performed using three previously developed deep learning models: Unet-R231, TotalSegmentator, MedSAM. Performance was assessed using quantitative metrics (volumetric similarity, Dice similarity coefficient, Hausdorff distance) and a qualitative measure (four-point clinical acceptability scale). Unet-R231 consistently outperformed TotalSegmentator and MedSAM in general, for different severity levels, and pathology categories (p<0.05). All models showed significant performance declines from mild to moderate-to-severe cases, particularly in volumetric similarity (p<0.05), without significant differences among lung sides or pathology types. Unet-R231 provided the most accurate automated lung segmentation among evaluated models with TotalSegmentator being a close second, though their performance declined significantly in moderate-to-severe cases, emphasizing the need for specialized model fine-tuning in severe pathology contexts.

肺分割AI医疗移植规划深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。