用分布建模提升肺癌病理切片肿瘤比例评分精度
Distribution-based deep multiple instance learning for tumor proportion scoring in NSCLC

- 基于分布的多实例学习,融合切片级标签预测整体表达分布
- 新方法在零值密集数据下准确率显著优于传统回归模型
- 适合病理图像分析、精准医疗领域研究者参考
非小细胞肺癌(NSCLC)中肿瘤比例评分(TPS)的准确评估对治疗方案制定和预后判断至关重要。主要挑战在于标注每张病理切片耗时费力,且具备认证资质的专家数量有限。多实例学习(MIL)已在切片级别预测TPS方面表现有效,但现有方法难以处理无表达(零类)图像。本文提出双模型架构:(1) 嵌入提取与多分类网络,捕捉单个组织切片块的组织学特征;(2) 基于嵌入聚合的MIL模型,用于预测表示整张切片TPS概率分布的零膨胀贝塔(ZIBeta)参数。仅使用切片级TPS标签作为监督信号,该端到端框架通过新型分布建模显著提升预测准确性和可解释性。实验表明,ZIBeta建模在性能上显著优于线性与岭回归基线,同时能有效反映预测置信度的分布集中性。
原文摘要 · Abstract (English)
Accurate assessment of tumor proportion score (TPS) in non-small cell lung cancer (NSCLC) is critical for treatment planning and prognosis. Key challenges include the tedious manual work required to annotate each slide, combined with the limited number of experts certified for this task. Multiple instance learning (MIL) has proven to be an effective approach for predicting TPS scores at the slide level; however, existing methods struggle with non-expressive (zero class) images. Our approach involves two models: (1) an embedding-extraction and multiclass-classification network that captures the histopathological features of individual patches, and (2) a MIL model that aggregates these embeddings to predict zero-inflated beta (ZIBeta) parameters representing the overall TPS probability distribution for the entire slide. Using only slide-level TPS scores as labels, we demonstrate how this end-to-end framework can leverage a novel distribution-based architecture to improve prediction accuracy and explainability. ZIBeta modeling significantly outperforms baseline linear and ridge regression while capturing expected accuracy through distribution concentration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。