基于组织学结构的切片方法提升癌症复发预测准确性和可解释性
Histology-informed tiling of whole tissue sections improves the interpretability and predictability of cancer relapse and genetic alterations
- 用语义分割提取腺体作为生物有意义的图像块
- 腺体级别分割精度Dice达0.83,显著提升基因变异检测AUC
- 识别出15个腺体簇,与复发、突变和病理分级相关
病理学家通过评估前列腺癌中的腺体等组织结构来确定癌症分级。然而,数字病理分析常采用忽略组织架构的网格切片方法,引入无关信息并降低可解释性。本文提出组织学引导切片(HIT),利用语义分割从全幻灯片图像(WSIs)中提取腺体,作为多实例学习(MIL)和表型分析的生物意义输入块。在ProMPT队列137例样本上训练,腺体级分割的Dice分数为0.83 ± 0.17。通过对ICGC-C和TCGA-PRAD队列共760张WSI提取38万枚腺体,HIT使检测与上皮-间质转化(EMT)及MYC相关拷贝数变异(CNVs)的MIL模型AUC提升10%,并发现15个腺体簇,其中多个与癌症复发、致癌突变及高格里森评分相关。HIT不仅提高了预测准确性与可解释性,还在特征提取阶段聚焦生物结构,优化计算效率。
原文摘要 · Abstract (English)
Histopathologists establish cancer grade by assessing histological structures, such as glands in prostate cancer. Yet, digital pathology pipelines often rely on grid-based tiling that ignores tissue architecture. This introduces irrelevant information and limits interpretability. We introduce histology-informed tiling (HIT), which uses semantic segmentation to extract glands from whole slide images (WSIs) as biologically meaningful input patches for multiple-instance learning (MIL) and phenotyping. Trained on 137 samples from the ProMPT cohort, HIT achieved a gland-level Dice score of 0.83 +/- 0.17. By extracting 380,000 glands from 760 WSIs across ICGC-C and TCGA-PRAD cohorts, HIT improved MIL models AUCs by 10% for detecting copy number variation (CNVs) in genes related to epithelial-mesenchymal transitions (EMT) and MYC, and revealed 15 gland clusters, several of which were associated with cancer relapse, oncogenic mutations, and high Gleason. Therefore, HIT improved the accuracy and interpretability of MIL predictions, while streamlining computations by focussing on biologically meaningful structures during feature extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。