提出SMILE方法,提升多中心肺腺癌空气传播病灶自动诊断准确率
SMILE: a Scale-aware Multiple Instance Learning Method for Multicenter STAS Lung Cancer Histopathology Diagnosis
- 引入尺度自适应注意力机制,缓解局部区域依赖
- 在三个数据集上分别识别出251和319个STAS样本,性能超临床平均水平
- 首个公开基准结果,助力计算病理学可解释性与临床落地
空气传播病灶(STAS)是肺癌中一种新发现的侵袭性病理特征,与不良预后相关且形态复杂。当前病理医生依赖耗时、主观性强的手动评估,亟需自动化精准诊断方案。本文整合来自多个中心的2,970张肺癌组织切片,重新诊断并公开发布三个STAS数据集:STAS CSU(医院)、STAS TCGA 和 STAS CPTAC,均包含病理特征诊断与临床数据。为应对STAS的偏差、稀疏性和异质性问题,提出尺度感知多实例学习(SMILE)方法。通过尺度自适应注意力机制,动态调整高关注实例,减少对局部区域的过度依赖,促进对STAS病灶的一致检测。大量实验表明,SMILE在STAS CSU数据集上成功诊断251例和319例样本,分别对应CPTAC与TCGA,AUC超越临床平均表现。11项开放基线结果首次建立,为未来计算病理技术的扩展、可解释性与临床集成奠定基础。数据集与代码已公开于https://anonymous.4open.science/r/IJCAI25-1DA1。
原文摘要 · Abstract (English)
Spread through air spaces (STAS) represents a newly identified aggressive pattern in lung cancer, which is known to be associated with adverse prognostic factors and complex pathological features. Pathologists currently rely on time consuming manual assessments, which are highly subjective and prone to variation. This highlights the urgent need for automated and precise diag nostic solutions. 2,970 lung cancer tissue slides are comprised from multiple centers, re-diagnosed them, and constructed and publicly released three lung cancer STAS datasets: STAS CSU (hospital), STAS TCGA, and STAS CPTAC. All STAS datasets provide corresponding pathological feature diagnoses and related clinical data. To address the bias, sparse and heterogeneous nature of STAS, we propose an scale-aware multiple instance learning(SMILE) method for STAS diagnosis of lung cancer. By introducing a scale-adaptive attention mechanism, the SMILE can adaptively adjust high attention instances, reducing over-reliance on local regions and promoting consistent detection of STAS lesions. Extensive experiments show that SMILE achieved competitive diagnostic results on STAS CSU, diagnosing 251 and 319 STAS samples in CPTAC andTCGA,respectively, surpassing clinical average AUC. The 11 open baseline results are the first to be established for STAS research, laying the foundation for the future expansion, interpretability, and clinical integration of computational pathology technologies. The datasets and code are available at https://anonymous.4open.science/r/IJCAI25-1DA1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。