arXiv:2606.11606cs.CV2026-06被引 1

冻结的视觉模型会丢失胸片中小病灶信号,但可通过局部区域恢复。

Frozen Foundation-Model Embeddings Discard Small-Lesion Signal in Chest Radiography: Implications for Pre-Deployment Evaluation

论文配图:Frozen Foundation-Model Embeddings Discard Small-Lesion Signal in Chest Radiography: Implications for Pre-Deployment Evaluation
图 1 · 摘自论文原文
  • 用局部区域池化从模型激活图中找回被丢弃的小病灶信息
  • 全局分类标记的准确率仅0.500-0.524,局部池化提升至0.899以上
  • 适用于医疗影像预部署评估,尤其关注小病灶检测的研究者

冻结的视觉变换器(ViT)基础模型嵌入正广泛用于下游胸片(CXR)分析流程,但其在前向传播中对小尺度、低对比度信号的保留或丢失情况尚未系统量化。本研究测试了五种冻结的ViT模型(RAD-DINO、DINOv2-B/14、DINOv3 ViT-7B、BiomedCLIP、MedSigLIP)和一种冻结的DINO预训练ResNet-50作为对照,在三个大规模胸片数据集(NIH-CXR14、MIMIC-CXR、Emory-CXR;总样本n=492,724)及ChestX-Det10(n=3,543;含1,462个小型病灶边界框,涵盖钙化、结节、肿块)上评估。通过小尺度扰动面板与区域感知边界框分层探针,比较三种聚合方式:分类标记(CLS)、所有最终层补丁平均(patch-mean)和边界框限制的局部补丁池化。在扰动测试中,CLS嵌入表现接近随机(AUC 0.500–0.524);patch-mean在等模糊和网状细粒度结构下与CLS无异,但在大方向模糊时优于后者;疾病任务整体AUC为0.642–0.913。而局部池化在同一前向传递中将AUC提升至接近1.0(各模型平均提升+0.412至+0.488),且ResNet-50控制组仍保持随机水平。在ChestX-Det10上,图像级分类存在小病灶与大病灶间的类别内差距最高达+0.243 AUC,而边界框级别局部池化在所有(模型×类别)组合中均恢复到AUC≥0.899。结果表明,冻结的ViT嵌入在全局聚合阶段悄然抑制了小尺度信号,但该信号可通过对特定感兴趣区域的补丁激活进行恢复。

原文摘要 · Abstract (English)

Frozen vision-transformer (ViT) foundation-model embeddings increasingly serve as the substrate for downstream chest-radiography (CXR) pipelines, yet where small-scale, low-contrast signal is retained or lost in the frozen forward pass has not been systematically quantified across architectures, pretraining domains, and objectives. We probed five frozen ViTs (RAD-DINO, DINOv2-B/14, DINOv3 ViT-7B, BiomedCLIP, MedSigLIP) and a frozen DINO-pretrained ResNet-50 architectural control across three large CXR cohorts (NIH-CXR14, MIMIC-CXR, Emory-CXR; aggregate pool n=492,724) and ChestX-Det10 (n=3,543; 1,462 small-lesion bounding boxes across Calcification, Nodule, Mass). Each model was evaluated with a small-scale-perturbation panel and a region-aware bounding-box-stratified probe on real lesions, comparing three pooling modes from the same forward pass: classification token (CLS), patch-mean (mean over all final-layer patch tokens), and bounding-box-restricted patch-local. On the perturbation panel, CLS embeddings sat at the chance floor (area under the ROC curve [AUC] 0.500-0.524); patch-mean was indistinguishable from CLS on iso-blur and reticular-fine cells but rose with CLS on larger directional-blur footprints, while disease AUC on globally decided tasks ranged 0.642-0.913. Patch-local probes recovered AUC ~1.0 from the same forward pass (per-model mean improvement +0.412 to +0.488); the ResNet-50 control reproduced the chance floor. On ChestX-Det10, image-level CLS classification showed within-class small-versus-large stratum gaps up to +0.243 AUC; bounding-box-level patch-local pooling on the same forward pass recovered AUC >= 0.899 on every (model x class) cell. Frozen ViT embeddings silently suppress small-scale signal at the global-aggregation step; the signal is recoverable from patch tokens conditional on a region of interest.

医学影像小病灶检测视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。