arXiv:2409.02143q-bio.GNcs.LG2024-09被引 32

MLOmics整合8314个癌种样本,专为机器学习模型训练与评估设计。

MLOmics: Cancer Multi-Omics Database for Machine Learning

  • 构建覆盖32种癌症的多组学数据集,支持机器学习直接使用
  • 包含8314名患者数据,涵盖四种组学类型和分层特征
  • 提供下游分析工具,适合生物信息与AI交叉研究者

将多种癌症研究视为机器学习问题,近年来在多组学分析与癌症研究中展现出显著潜力。推动这些成功模型发展的关键在于高质量、数据量充足且预处理完善的训练数据集。然而,尽管已有多个公开数据门户(如TCGA多组学计划或LinkedOmics等开放数据库),这些数据集仍不适用于现有机器学习模型的即插即用。本文提出MLOmics,一个开源的癌症多组学数据库,旨在更好支持生物信息学与机器学习模型的开发与评估。MLOmics包含8,314例患者样本,覆盖全部32种癌症类型,涵盖四种组学类型、分层特征及广泛基准测试。同时提供下游分析支持与生物知识关联功能,促进跨学科研究。

原文摘要 · Abstract (English)

Framing the investigation of diverse cancers as a machine learning problem has recently shown significant potential in multi-omics analysis and cancer research. Empowering these successful machine learning models are the high-quality training datasets with sufficient data volume and adequate preprocessing. However, while there exist several public data portals, including The Cancer Genome Atlas (TCGA) multi-omics initiative or open-bases such as the LinkedOmics, these databases are not off-the-shelf for existing machine learning models. In this paper, we introduce MLOmics, an open cancer multi-omics database aiming at serving better the development and evaluation of bioinformatics and machine learning models. MLOmics contains 8,314 patient samples covering all 32 cancer types with four omics types, stratified features, and extensive baselines. Complementary support for downstream analysis and bio-knowledge linking are also included to support interdisciplinary analysis.

癌症研究多组学机器学习数据库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。