医学图像中冻结视觉模型的抗噪训练,需根据噪声类型选方法,不能一概而论。
Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure

- 构建跨数据集基准,测试8种抗噪方法在5个医疗数据集上的表现
- 小损失假设失效:在40%异构噪声下,准确率差距达18.8个百分点
- 提出特征空间选择器,帮助实际应用中避免错误方法选择
冻结视觉基础模型(VFMs)因其轻量分类头,在医学影像中广泛应用,但针对该场景的抗噪学习方法仍不清晰,多数依赖于端到端训练继承的小损失假设。本文在五个医学数据集、三种骨干网络、两种噪声类型、五种噪声率(共150组条件,6000次训练)下,系统评估了八种抗噪方法,以平衡准确率为指标。结果显示无通用最优方法:弗里德曼排名显示显著差异(χ² = 333.2,p = 4.77 × 10⁻⁶⁸),ELR在49/150条件下胜出,CUFIT平均排名最佳(2.51)。方法选择的代价随噪声加剧急剧上升,从清洁数据下的4.5个百分点增至异构40%噪声下的18.8个百分点。深入分析发现,在冻结DINOv2特征下,干净与噪声样本损失分布重叠率达53–61%,且在异构噪声中预测一致性的稳定性远超损失排序(精度下降仅3pp vs. 13pp)。在ISIC2019数据集上,40%异构噪声下Co-Teaching整体准确率达68%,但平衡准确率降至35.1%,三类少数类召回率为零。研究表明,抗噪训练应作为情境感知的方法选择问题,而非寻找单一优解。最后提供基于证据的指导和低后悔特征空间选择器用于实践推荐。
原文摘要 · Abstract (English)
Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in medical imaging because they offer efficient and reproducible deployment. Yet noisy-label learning methods for this frozen-feature regime remain poorly understood, and most existing methods still rely on a small-loss assumption inherited from end-to-end training. We present a controlled benchmark of eight noisy-label methods across five medical datasets, three backbones, two noise types, and five noise rates (150 conditions, 6,000 training runs), evaluated with balanced accuracy. The benchmark shows that there is no universal winner: Friedman ranking over the 150 conditions yields $χ^2 = 333.2$ ($p = 4.77 \times 10^{-68}$), ELR wins the most conditions (49/150), while CUFIT attains the best mean rank (2.51). The practical cost of method choice grows sharply with noise severity, from 4.5pp on clean data to 18.8pp at asymmetric 40\% noise. To explain these benchmark-level patterns, we revisit the small-loss assumption in a representative high-risk regime. Under frozen DINOv2 features, clean and noisy loss distributions overlap by 53--61\%, and matched-rate clean-sample detection shows that prediction agreement is markedly more stable than loss ranking under asymmetric noise (3pp vs.\ 13pp precision drop). On ISIC2019 with asymmetric 40\% noise, Co-Teaching reaches 68\% overall accuracy while collapsing to 35.1\% balanced accuracy with zero recall on three minority classes. Together, these results recast noisy-label learning for frozen VFMs as a regime-aware method-selection problem rather than a search for a single dominant algorithm. We conclude with evidence-based guidance and a low-regret feature-space selector for practical recommendation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。