用多数据源预训练模型,提升前列腺癌分级准确率
ProsMAE: Multi-Source MAE Pretraining for ISUP Grade Classification

- 融合三个数据集的组织切片,用掩码自编码器学习病理表征
- 在PANDA数据集上比基线模型平均加权肯德尔系数提升1.8%
- 适合做病理图像分类的迁移学习,尤其关注前列腺癌分级
全切片图像(WSIs)为计算病理学提供丰富诊断信息,但其吉字节级规模、染色差异、扫描仪差异、组织伪影以及专家标注有限,使得模型训练困难。本文提出一种多源掩码自编码器(MAE)框架ProsMAE,用于组织病理学表征学习。使用前列腺癌分级评估(PANDA)、淋巴结转移癌挑战2017(CAMELYON17)和乳腺癌亚型分类(BRACS)的数据进行预训练,使编码器接触多样化的组织形态与采集条件。通过冻结编码器并添加线性分类头,在国际泌尿病理学会(ISUP)分级任务中进行迁移学习(ProsCLS)。ProsMAE在评估的不相交PANDA划分下,平均验证加权肯德尔系数(QWK)高于原始MAE冻结线性探测基线。重复分组评估仍需进行,以进一步验证跨划分组合的鲁棒性。
原文摘要 · Abstract (English)
Whole slide images (WSIs) provide rich diagnostic information for computational pathology, but their gigapixel scale, stain variation, scanner differences, tissue artifacts, and limited expert annotation make robust model training challenging. This paper presents a multi-source Masked Autoencoder (MAE) framework, named ProsMAE, for histopathology representation learning. Tiles from Prostate cANcer graDe Assessment (PANDA), CAncer MEtastases in LYmph nOdes challeNge 2017 (CAMELYON17), and BReAst Carcinoma Subtyping (BRACS) are used for ProsMAE pretraining to expose the encoder to diverse tissue morphology and acquisition conditions. The learned encoder is transferred for International Society of Urological Pathology (ISUP) grade classification through ProsCLS, using a frozen encoder and a linear classification head. ProsMAE achieved a higher mean validation quadratic weighted kappa (QWK) than the vanilla MAE frozen linear-probe baseline under the evaluated disjoint PANDA split. Repeated-split evaluation remains necessary to further establish robustness across split compositions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。