用H&E引导学习SIM图像结构,提升病理成像的跨模态泛化能力。
SIMPLER: H&E-Informed Representation Learning for Structured Illumination Microscopy

- 以H&E为语义锚点,通过对抗、对比和重建三重目标对齐SIM与H&E特征
- 在多个下游任务中超越从零训练和仅用H&E预训练的模型表现
- 适合需要快速无损成像的术中诊断与点护理场景
结构光照明显微镜(SIM)可在不染色或切片的情况下实现新鲜组织的快速高对比度光学切片,适用于术中和点护理诊断。当前数字病理学中的大模型多基于薄组织切片的H&E和免疫组化(IHC)数据训练,但未针对厚组织荧光模态如SIM进行优化。直接迁移至SIM时,因模态差异性能受限,简单微调常导致对特定外观过拟合而非捕捉组织结构。本文提出SIMPLER(Structured Illumination Microscopy-Powered Learning for Embedding Representations),一种跨模态自监督预训练框架,利用H&E作为语义锚点,学习可复用的SIM表征。H&E包含与临床标注一致的细胞与腺体结构信息,而SIM提供快速非破坏性成像。预训练阶段通过对抗、对比和重构目标逐步对齐SIM与H&E特征,促使SIM嵌入内化组织结构,同时保留模态特异性。单一预训练编码器在多实例学习和形态聚类等任务上持续优于从零训练或仅用H&E预训练的模型。结果表明,基于组织学引导的跨模态预训练可生成生物学合理的SIM嵌入,适用于广泛下游应用。
原文摘要 · Abstract (English)
Structured Illumination Microscopy (SIM) enables rapid, high-contrast optical sectioning of fresh tissue without staining or physical sectioning, making it promising for intraoperative and point-of-care diagnostics. Recent foundation and large-scale self-supervised models in digital pathology have demonstrated strong performance on section-based modalities such as Hematoxylin and Eosin (H&E) and immunohistochemistry (IHC). However, these approaches are predominantly trained on thin tissue sections and do not explicitly address thick-tissue fluorescence modalities such as SIM. When transferred directly to SIM, performance is constrained by substantial modality shift, and naive fine-tuning often overfits to modality-specific appearance rather than underlying histological structure. We introduce SIMPLER (Structured Illumination Microscopy-Powered Learning for Embedding Representations), a cross-modality self-supervised pretraining framework that leverages H&E as a semantic anchor to learn reusable SIM representations. H&E encodes rich cellular and glandular structure aligned with established clinical annotations, while SIM provides rapid, nondestructive imaging of fresh tissue. During pretraining, SIM and H&E are progressively aligned through adversarial, contrastive, and reconstruction-based objectives, encouraging SIM embeddings to internalize histological structure from H&E without collapsing modality-specific characteristics. A single pretrained SIMPLER encoder transfers across multiple downstream tasks, including multiple instance learning and morphological clustering, consistently outperforming SIM models trained from scratch or H&E-only pretraining. These results suggest that histology-guided cross-modal pretraining yields biologically grounded SIM embeddings suitable for broad downstream reuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。