将基因互作先验知识注入深度学习,提升从病理切片预测基因表达的准确率
Prior Knowledge Injection into Deep Learning Models Predicting Gene Expression from Whole Slide Images
- 通过通用框架将基因互作知识融入深度模型,增强预测鲁棒性
- 乳腺癌实验中平均多识别983个显著基因,14项在独立数据集上有效
- 适用于多种模型架构,适合病理与基因组联合研究者
癌症诊断和预后主要依赖年龄、肿瘤分级等临床参数,日益补充基因表达等分子数据。但基因测序成本高且延迟诊疗流程。深度学习可从全幻灯片图像(WSIs)的形态特征预测分子信息,提供经济高效的分子标记替代方案。然而现有方法仍缺乏足够稳健性以完全替代测序。本文提出一种模型无关框架,将基因-基因互作先验知识注入深度学习架构,提升预测准确性和鲁棒性。该框架设计通用,可灵活适配多种模型。在乳腺癌案例研究中,该策略在全部18组实验中平均增加983个显著基因(共25,761个),其中14组在独立数据集上仍有效。结果表明,注入先验知识可显著提升从WSIs预测基因表达性能,适用于广泛模型架构。
原文摘要 · Abstract (English)
Cancer diagnosis and prognosis primarily depend on clinical parameters such as age and tumor grade, and are increasingly complemented by molecular data, such as gene expression, from tumor sequencing. However, sequencing is costly and delays oncology workflows. Recent advances in Deep Learning allow to predict molecular information from morphological features within Whole Slide Images (WSIs), offering a cost-effective proxy of the molecular markers. While promising, current methods lack the robustness to fully replace direct sequencing. Here we aim to improve existing methods by introducing a model-agnostic framework that allows to inject prior knowledge on gene-gene interactions into Deep Learning architectures, thereby increasing accuracy and robustness. We design the framework to be generic and flexibly adaptable to a wide range of architectures. In a case study on breast cancer, our strategy leads to an average increase of 983 significant genes (out of 25,761) across all 18 experiments, with 14 generalizing to an increase on an independent dataset. Our findings reveal a high potential for injection of prior knowledge to increase gene expression prediction performance from WSIs across a wide range of architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。