arXiv:2603.04081cs.CVq-bio.QM2026-03

小图像块下,专用模型比大模型更高效准确

Revisiting the Role of Foundation Models in Cell-Level Histopathological Image Analysis under Small-Patch Constraints -- Effects of Training Data Scale and Blur Perturbations on CNNs and Vision Transformers

  • 用40x40小图训练专用模型,数据越多越准
  • 自定义ViT在小图上表现最优,推理成本低
  • 大模型在小图上难提升,抗模糊能力无优势

背景与目标:细胞级病理图像分析需处理极小图像块(40x40像素),远低于标准ImageNet分辨率。当前尚不明确现代深度学习架构和基础模型在此约束下能否学习到鲁棒且可扩展的表示。本文系统评估了小图块细胞分类任务中不同架构的适用性及数据规模的影响。方法:基于303例结直肠癌样本(CD103/CD8免疫染色),生成185,432个标注细胞图像。训练了8种特定任务模型,覆盖256至16,384样本/类的不同数据规模;评估了3种基础模型在输入重采样至224x224后的线性探针与微调表现,并通过高斯模糊扰动评估其鲁棒性。结果:任务专用模型随数据量增加持续提升,而基础模型在中等样本量后趋于饱和。专为小图优化的Vision Transformer(CustomViT)准确率最高,显著优于所有基础模型且推理开销更低。各架构在模糊扰动下的鲁棒性相近,未见基础模型有明显优势。结论:在极端空间约束下,一旦具备足够训练数据,专用模型比基础模型更有效、更高效。更高准确率不等于更强鲁棒性,大预训练模型在小图场景中收益有限。

原文摘要 · Abstract (English)

Background and objective: Cell-level pathological image analysis requires working with extremely small image patches (40x40 pixels), far below standard ImageNet resolutions. It remains unclear whether modern deep learning architectures and foundation models can learn robust and scalable representations under this constraint. We systematically evaluated architectural suitability and data-scale effects for small-patch cell classification. Methods: We analyzed 303 colorectal cancer specimens with CD103/CD8 immunostaining, generating 185,432 annotated cell images. Eight task-specific architectures were trained from scratch at multiple data scales (FlagLimit: 256--16,384 samples per class), and three foundation models were evaluated via linear probing and fine-tuning after resizing inputs to 224x224 pixels. Robustness to blur was assessed using pre- and post-resize Gaussian perturbations. Results: Task-specific models improved consistently with increasing data scale, whereas foundation models saturated at moderate sample sizes. A Vision Transformer optimized for small patches (CustomViT) achieved the highest accuracy, outperforming all foundation models with substantially lower inference cost. Blur robustness was comparable across architectures, with no qualitative advantage observed for foundation models. Conclusion: For cell-level classification under extreme spatial constraints, task-specific architectures are more effective and efficient than foundation models once sufficient training data are available. Higher clean accuracy does not imply superior robustness, and large pre-trained models offer limited benefit in the small-patch regime.

病理图像小图分析视觉变换器模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。