对比多种基因组模型,发现主成分分析仍是最优扰动分析工具
Benchmarking Transcriptomics Foundation Models for Perturbation Analysis : one PCA still rules them all
- 构建生物可解释的评估框架,系统比较预训练模型性能
- scVI与PCA在真实场景中显著优于现有基础模型
- 适用于生物医学研究者评估基因表达模型可靠性
理解基因、化合物及其相互作用在生物体中的关系仍受限于技术瓶颈和数据复杂性。深度学习在多类型数据中探索这些关系展现出潜力,但转录组学因高噪声和数据稀缺而未被充分应用。近年来测序技术进步为挖掘有价值信息提供了新机遇,尤其是众多转录组学基础模型的兴起,然而尚无稳健基准来评估其在扰动分析中的有效性。本文提出一个生物学动机驱动的评估框架,构建了扰动分析任务层级,用于比较预训练基础模型与经典学习方法在转录组数据上的表现。我们整合了来自不同测序技术和细胞系的多样化公开数据集以评估模型性能。结果表明,scVI和主成分分析(PCA)在理解生物扰动方面远优于现有基础模型,尤其在真实应用场景中表现突出。
原文摘要 · Abstract (English)
Understanding the relationships among genes, compounds, and their interactions in living organisms remains limited due to technological constraints and the complexity of biological data. Deep learning has shown promise in exploring these relationships using various data types. However, transcriptomics, which provides detailed insights into cellular states, is still underused due to its high noise levels and limited data availability. Recent advancements in transcriptomics sequencing provide new opportunities to uncover valuable insights, especially with the rise of many new foundation models for transcriptomics, yet no benchmark has been made to robustly evaluate the effectiveness of these rising models for perturbation analysis. This article presents a novel biologically motivated evaluation framework and a hierarchy of perturbation analysis tasks for comparing the performance of pretrained foundation models to each other and to more classical techniques of learning from transcriptomics data. We compile diverse public datasets from different sequencing techniques and cell lines to assess models performance. Our approach identifies scVI and PCA to be far better suited models for understanding biological perturbations in comparison to existing foundation models, especially in their application in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。