arXiv:2512.24294cs.CVcs.AI2025-12被引 1

针对肺癌筛查CT图像,提出一套可量化质量控制流程,提升大模型预测性能。

Virtual-Eyes: Quantitative Validation of a Lung CT Quality-Control Pipeline for Foundation-Model Cancer Risk Prediction

  • 设计16位CT图像质量控制流程,强制统一分辨率并剔除无效数据
  • 使通用大模型在肺结节预测上的患者级准确率从61.9%提升至73.5%
  • 揭示通用模型与专用模型对预处理的差异化响应,指导临床应用选择

深度学习用于低剂量CT肺癌筛查时,稳健的预处理常缺乏量化评估。本文开发并验证了虚拟眼(Virtual-Eyes)这一临床驱动的16位CT质量控制流程,测量其对通用基础模型与专用模型的影响。该流程强制512x512像素分辨率,剔除短序列或非诊断性扫描,利用亨氏单位滤波和双肺覆盖率评分提取连续肺块,同时保留原始16位数据结构。基于NLST中765名患者(182例癌症,583例非癌症),我们使用冻结编码器计算RAD-DINO与Merlin的切片级嵌入,并训练无泄漏的患者级MLP头;同时在原始与虚拟眼输入下评估Sybil与2D ResNet-18基线模型,不进行主干重训练。结果表明,虚拟眼将RAD-DINO切片级AUC从0.576提升至0.610,患者级AUC从0.646升至0.683(均值池化),以及从0.619升至0.735(最大池化),且校准性能显著改善(布里尔分数从0.188降至0.112)。相反,Sybil与ResNet-18在虚拟眼处理后性能下降(Sybil AUC从0.886降至0.837;ResNet-18从0.571降至0.596),显示上下文依赖性与捷径学习现象;Merlin则无论预处理方式,迁移能力有限(AUC约0.507至0.567)。结果表明,解剖目标导向的质量控制可稳定并提升通用基础模型工作流,但可能破坏已适应原始临床环境的专用模型。

原文摘要 · Abstract (English)

Robust preprocessing is rarely quantified in deep-learning pipelines for low-dose CT (LDCT) lung cancer screening. We develop and validate Virtual-Eyes, a clinically motivated 16-bit CT quality-control pipeline, and measure its differential impact on generalist foundation models versus specialist models. Virtual-Eyes enforces strict 512x512 in-plane resolution, rejects short or non-diagnostic series, and extracts a contiguous lung block using Hounsfield-unit filtering and bilateral lung-coverage scoring while preserving the native 16-bit grid. Using 765 NLST patients (182 cancer, 583 non-cancer), we compute slice-level embeddings from RAD-DINO and Merlin with frozen encoders and train leakage-free patient-level MLP heads; we also evaluate Sybil and a 2D ResNet-18 baseline under Raw versus Virtual-Eyes inputs without backbone retraining. Virtual-Eyes improves RAD-DINO slice-level AUC from 0.576 to 0.610 and patient-level AUC from 0.646 to 0.683 (mean pooling) and from 0.619 to 0.735 (max pooling), with improved calibration (Brier score 0.188 to 0.112). In contrast, Sybil and ResNet-18 degrade under Virtual-Eyes (Sybil AUC 0.886 to 0.837; ResNet-18 AUC 0.571 to 0.596) with evidence of context dependence and shortcut learning, and Merlin shows limited transferability (AUC approximately 0.507 to 0.567) regardless of preprocessing. These results demonstrate that anatomically targeted QC can stabilize and improve generalist foundation-model workflows but may disrupt specialist models adapted to raw clinical context.

CT质量控制肺癌筛查基础模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。