arXiv:2606.06950cs.CVcs.AI2026-06

2.5D CNN在肺部CT分类中平衡性能与稳定性最佳

When is 3D Worth It? A Resource-Performance Frontier for CNNs and Transformers in Lung CT

  • 比较2D、2.5D、3D输入对CNN与Transformer的影响
  • 2.5D CNN ROC-AUC达0.682,稳定性最优
  • 3D模型易出现阈值不稳,Transformer会全正预测

三维模型常被视为体积医学影像的首选,但其实际价值取决于性能提升是否足以抵消计算成本与复杂度。本文不提出新架构,而是在固定训练协议下,研究输入维度(2D、2.5D、3D)对卷积神经网络(CNN)和视觉变换器(ViT)行为的影响。基于无泄漏的NLST队列(n=1,977)及辅助的LIDC-IDRI数据,发现2.5D CNN在区分度与稳定性之间具有最优权衡(ROC-AUC 0.682,95% CI [0.546, 0.799]),且运行点稳定。相反,3D CNN表现出阈值不稳,而Transformer则出现退化预测(如全为阳性)。置信区间宽且重叠,因此将结果呈现为受控资源-性能边界与失效模式分类,而非确定性优势声明。对于类别不平衡的肺癌筛查分类任务,2D与2.5D输入在性能、稳定性与计算效率间提供了更可靠的折衷方案。

原文摘要 · Abstract (English)

Three-dimensional models are widely assumed preferable for volumetric medical imaging, yet their practical value depends on whether performance gains justify added computational cost and complexity. Rather than proposing a new architecture, we study how input dimensionality (2D, 2.5D, 3D) affects model behavior across convolutional neural networks (CNNs) and Vision Transformers (ViTs) under a fixed training protocol. Using a leakage-free NLST cohort (n = 1,977) with supporting LIDC-IDRI data, we find that the 2.5D CNN offers the most favorable discrimination-stability trade-off in our comparison (ROC-AUC 0.682, 95% CI [0.546, 0.799]) with a stable operating point. In contrast, 3D CNNs show threshold instability, and transformers exhibit degenerate predictions, such as all-positive predictions. Confidence intervals are wide and overlapping, so we present these results as a controlled resource-performance frontier and a failure-mode taxonomy rather than as definitive superiority claims. For class-imbalanced lung cancer screening classification, 2D and 2.5D inputs provide a more reliable trade-off between performance, stability, and computational efficiency than full 3D representations.

肺部CT2.5DCNNTransformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。