用Vision Transformer从胸片预测1-2年后的肺癌,提升早期筛查潜力。
Extending the Horizon of Early Diagnosis: Lung Cancer Prediction with Vision Transformers

- 基于ViT模型分析海量胸片,通过迁移学习与采样优化应对数据不平衡
- 预训练模型比从零训练提升6-10个百分点的AUC,最高敏感性达85%以上
- 适合医疗AI研究者、放射科医生及早期筛查系统开发者参考
肺癌仍是全球癌症致死主因,早期诊断对提高生存率至关重要。然而,早期恶性病变在胸片上常表现微妙,给放射科医生带来挑战。本研究评估视觉变换器(ViTs)在临床诊断前1至2年预测肺癌的可行性。分析了来自波士顿雅马哈平原退伍军人医院的91,020次影像检查中的259,361张胸片,数据存在约1:150的极端类别不平衡问题,通过混合欠采样与过采样及加权损失优化解决。评估了三种ViT配置:从零训练模型、ImageNet预训练模型、Corona预训练模型微调。迁移学习显著提升性能,预训练模型相比基线在AUC上提升6-10个百分点,平衡准确率提升约10-12%。ImageNet预训练模型整体表现最稳定,Corona预训练模型在某些场景下敏感性更高但波动较大。中等采样比例(如1:1欠采样与1.5:2过采样)在敏感性、精确率与计算效率间取得良好权衡,运行时间最多减少70%而性能损失较小。结果表明ViTs具备从常规胸片中预测早期肺癌风险的潜力。尽管当前性能尚未达到临床部署标准,但仍支持进一步开发基于ViT的分诊系统,用于筛选高危人群进行早期评估。
原文摘要 · Abstract (English)
Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage malignancies can be subtle on chest X-rays, creating challenges for radiologists. This study evaluates Vision Transformers (ViTs) for predicting lung cancer one to two years before clinical diagnosis. We analyzed 259,361 chest X-rays from 91,020 imaging studies at the Jamaica Plains VA Hospital in Boston, MA. The dataset showed extreme class imbalance, approximately 1:150 cancer to non-cancer, which was addressed using hybrid under- and over-sampling and class-weighted loss optimization. Three ViT configurations were evaluated: a model trained from scratch, an ImageNet-pretrained model, and a Corona-pretrained model fine-tuned on the lung cancer dataset. Transfer learning improved performance, with pretrained models exceeding the scratch baseline by 6-10 percentage points in AUC and about 10-12 percent in balanced accuracy. ImageNet-pretrained models showed the most stable overall performance, while Corona-pretrained models achieved higher sensitivity in some settings but greater variability. Moderate resampling ratios, including 1:1 undersampling and 1.5:2 oversampling, provided favorable trade-offs between sensitivity, precision, and computational efficiency, reducing runtime by up to 70 percent without major performance loss. These findings demonstrate the potential of ViTs for early lung cancer risk prediction from routine chest X-rays. Although performance remains below clinical deployment thresholds, the results support further development of ViT-based triage systems to flag high-risk patients for earlier evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。