用卷积网络做病理图像生成预训练,提升细粒度细胞预测效果。
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction

- 采用全卷积结构在像素空间进行掩码扩散预训练
- 在有限标注下仍优于主流ViT模型,小参数微调即达领先性能
- 适合需要高精度细胞级分析的病理图像任务
细胞级密集预测是计算病理学的核心挑战,受限于细微组织结构、强域偏移及昂贵的密集标注。现有基于ViT的病理基础模型依赖块标记化,破坏空间连续性,削弱局部形态细节。为此,我们提出卷积型掩码扩散基础模型(CMD),采用全卷积的ConvNeXt-UNet骨干,在像素空间进行自监督生成预训练,并通过自适应归一化融合冻结的病理基础模型特征。实验表明,CMD在多个病理密集预测任务中持续优于现有ViT模型,甚至超越当前最优端到端分割方法,且仅需微调少量任务特定参数。在标注有限条件下优势更显著,展现出更强鲁棒性与泛化能力。结果表明,纯卷积架构亦可作为高效的病理基础模型,在当前以ViT为主导的范式中实现领先性能,为细粒度病理理解提供可扩展、高性能的解决方案,更好保留组织学结构先验。
原文摘要 · Abstract (English)
Cell-level dense prediction is central to computational pathology, but remains challenging due to fine-grained histological structures, strong domain shifts, and costly dense annotations. Existing ViT-based pathology foundation models rely on patch tokenization, which can disrupt spatial continuity and weaken local morphological details needed for cell-level prediction. To address this, we propose Masked-Diffusion Convolutional Foundation Models, termed ConvNeXt Masked-Diffusion (CMD), a self-supervised convolutional generative pretraining framework for dense pathology representation learning. CMD uses a fully convolutional ConvNeXt-UNet backbone, performs masked-diffusion pretraining in pixel space, and incorporates frozen pathology foundation model features through adaptive normalization. Experimental results demonstrate that CMD consistently outperforms existing ViT-based pathology foundation models and even surpasses state-of-the-art end-to-end segmentation methods while fine-tuning only a small number of task-specific parameters across multiple pathology dense prediction tasks. The advantage is particularly pronounced under limited annotation settings, where CMD exhibits stronger robustness and generalization ability. Our findings suggest that purely convolutional architectures can also serve as competitive pathology foundation models for cell-level dense prediction, achieving leading performance within the current ViT-dominated paradigm and providing a scalable, high-performance solution that better preserves histological structural priors for fine-grained pathology understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。