对比肺结节与全胸影像模型,发现聚焦病灶的AI更准。
Refining Focus in AI for Lung Cancer: Comparing Lesion-Centric and Chest-Region Models with Performance Insights from Internal and External Validation
- 用病灶局部图像训练模型,比用整幅胸腔图像更有效。
- 外部数据集上病灶模型AUC达0.90,优于全胸模型的0.81。
- 对腺癌和不同厂家CT有更强适应性,适合精准诊断场景。
背景:基于AI的分类模型对提升肺癌诊断至关重要,但病灶级与胸腔区域模型在内部与外部数据集上的表现差异尚不明确。目的:评估病灶级与胸腔区域模型在肺癌分类中的表现,比较其在内部杜克肺结节数据集2024(DLND24)及外部(LUNA16、NLST)数据集上的效果,并针对人群特征、组织学类型和影像特征进行亚组分析。材料与方法:训练两个AI模型,一个使用病灶中心块(64,64,64),另一个使用胸腔区域块(512,512,8)。内部验证使用DLND24,外部验证采用LUNA16和NLST数据集。通过AUC-ROC评估性能,并使用DeLong检验进行统计比较,辅以梯度可视化与概率分布分析。结果:病灶级模型在各数据集上均表现更优。内部验证中,病灶模型AUC为0.71(CI: 0.61–0.81),胸腔模型为0.68(0.57–0.77)。外部验证显示类似趋势:在LUNA16上,病灶模型AUC为0.90(0.87–0.92),胸腔模型为0.81(0.79–0.82);在NLST上分别为0.81和0.77。亚组分析显示,病灶模型在腺癌及特定影像条件(如不同CT制造商)下优势显著。结论:病灶级模型在外部数据集及复杂亚组中表现更佳,提示其在精准肺癌诊断中具有临床应用潜力。
原文摘要 · Abstract (English)
Background: AI-based classification models are essential for improving lung cancer diagnosis. However, the relative performance of lesion-level versus chest-region models in internal and external datasets remains unclear. Purpose: This study evaluates the performance of lesion-level and chest-region models for lung cancer classification, comparing their effectiveness across internal Duke Lung Nodule Dataset 2024 (DLND24) and external (LUNA16, NLST) datasets, with a focus on subgroup analyses by demographics, histology, and imaging characteristics. Materials and Methods: Two AI models were trained: one using lesion-centric patches (64,64,64) and the other using chest-region patches (512,512,8). Internal validation was conducted on DLND24, while external validation utilized LUNA16 and NLST datasets. The models performances were assessed using AUC-ROC, with subgroup analyses for demographic, clinical, and imaging factors. Statistical comparisons were performed using DeLongs test. Gradient-based visualizations and probability distribution were further used for analysis. Results: The lesion-level model consistently outperformed the chest-region model across datasets. In internal validation, the lesion-level model achieved an AUC of 0.71(CI: 0.61-0.81), compared to 0.68(0.57-0.77) for the chest-region model. External validation showed similar trends, with AUCs of 0.90(0.87-0.92) and 0.81(0.79-0.82) on LUNA16 and NLST, respectively. Subgroup analyses revealed significant advantages for lesion-level models in certain histological subtypes (adenocarcinoma) and imaging conditions (CT manufacturers). Conclusion: Lesion-level models demonstrate superior classification performance, especially for external datasets and challenging subgroups, suggesting their clinical utility for precision lung cancer diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。