用病理切片预测肺癌基因突变,准确率达96.5%。
Predicting EGFR Mutation in LUAD from Histopathological Whole-Slide Images Using Pretrained Foundation Model and Transfer Learning: An Indian Cohort Study
- 基于视觉变换器和注意力机制的深度学习框架
- 在印度队列和外部数据集上AUC分别达0.933和0.965
- 小样本训练下表现优于以往研究,适合资源有限地区
肺腺癌(LUAD)是非小细胞肺癌的亚型,其中约46%携带EGFR基因突变。携带该突变的患者可使用特定酪氨酸激酶抑制剂(TKIs)治疗,因此预测突变状态对临床决策至关重要。H&E染色全切片图像(WSI)是癌症分期与分型的常规筛查手段,尤其在东南亚人群中突变率显著高于高加索人群(39-64% vs 7-22%)。本研究提出一种基于视觉变换器(ViT)病理基础模型与注意力多实例学习(ABMIL)架构的深度学习框架,从H&E WSI中预测EGFR突变状态。模型在印度队列(170张WSI)上训练,并在两个独立测试集上评估:内部测试集(30张来自印度队列)和外部TCGA测试集(86张WSI)。模型在两组数据上均表现一致,内部测试集AUC为0.933(±0.010),外部测试集AUC达0.965(±0.015)。该框架可在小样本数据上高效训练,性能优于多项先前研究,无论训练域如何。研究表明,利用基础模型与注意力机制,可准确预测常规病理切片中的EGFR突变状态,尤其适用于资源有限环境。
原文摘要 · Abstract (English)
Lung adenocarcinoma (LUAD) is a subtype of non-small cell lung cancer (NSCLC). LUAD with mutation in the EGFR gene accounts for approximately 46% of LUAD cases. Patients carrying EGFR mutations can be treated with specific tyrosine kinase inhibitors (TKIs). Hence, predicting EGFR mutation status can help in clinical decision making. H&E-stained whole slide imaging (WSI) is a routinely performed screening procedure for cancer staging and subtyping, especially affecting the Southeast Asian populations with significantly higher incidence of the mutation when compared to Caucasians (39-64% vs 7-22%). Recent progress in AI models has shown promising results in cancer detection and classification. In this study, we propose a deep learning (DL) framework built on vision transformers (ViT) based pathology foundation model and attention-based multiple instance learning (ABMIL) architecture to predict EGFR mutation status from H&E WSI. The developed pipeline was trained using data from an Indian cohort (170 WSI) and evaluated across two independent datasets: Internal test (30 WSI from Indian cohort) set, and an external test set from TCGA (86 WSI). The model shows consistent performance across both datasets, with AUCs of 0.933 (+/-0.010), and 0.965 (+/-0.015) for the internal and external test sets respectively. This proposed framework can be efficiently trained on small datasets, achieving superior performance as compared to several prior studies irrespective of training domain. The current study demonstrates the feasibility of accurately predicting EGFR mutation status using routine pathology slides, particularly in resource-limited settings using foundation models and attention-based multiple instance learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。