首个综合多组学评估基准,助力生物模型选型与分析
COMET: Benchmark for Comprehensive Biological Multi-omics Evaluation Tasks and Language Models
- 构建涵盖单组学、跨组学与多组学的多样化任务与数据集
- 首次系统评估DNA/RNA/蛋白语言模型及新提出的多组学方法性能
- 适合生物信息学、AI制药及多模态模型研究者使用
DNA、RNA和蛋白质作为中心法则的关键组成部分,在维持生命过程中保障基因表达与执行的准确性。尽管这些分子的研究已深刻影响医学、农业与工业,但机器学习方法从传统统计到深度学习、大语言模型的多样性,使研究人员在选择适用于特定任务(尤其是跨组学与多组学)的模型时面临挑战,主要源于缺乏全面的评估基准。为此,我们提出了首个综合性多组学评估基准COMET(Benchmark for Biological COmprehensive Multi-omics Evaluation Tasks and Language Models),旨在评估模型在单组学、跨组学与多组学任务中的表现。首先,我们整理并开发了覆盖DNA、RNA和蛋白质关键结构与功能方面的多样化下游任务与数据集,包括跨越多个组学层次的任务。其次,我们评估了现有的DNA、RNA、蛋白质基础语言模型以及新提出的多组学方法,揭示其在整合与分析不同生物模态数据方面的性能表现。该基准旨在识别多组学研究中的关键问题,指引未来方向,推动通过整合多组学数据分析深化对生物过程的理解。
原文摘要 · Abstract (English)
As key elements within the central dogma, DNA, RNA, and proteins play crucial roles in maintaining life by guaranteeing accurate genetic expression and implementation. Although research on these molecules has profoundly impacted fields like medicine, agriculture, and industry, the diversity of machine learning approaches-from traditional statistical methods to deep learning models and large language models-poses challenges for researchers in choosing the most suitable models for specific tasks, especially for cross-omics and multi-omics tasks due to the lack of comprehensive benchmarks. To address this, we introduce the first comprehensive multi-omics benchmark COMET (Benchmark for Biological COmprehensive Multi-omics Evaluation Tasks and Language Models), designed to evaluate models across single-omics, cross-omics, and multi-omics tasks. First, we curate and develop a diverse collection of downstream tasks and datasets covering key structural and functional aspects in DNA, RNA, and proteins, including tasks that span multiple omics levels. Then, we evaluate existing foundational language models for DNA, RNA, and proteins, as well as the newly proposed multi-omics method, offering valuable insights into their performance in integrating and analyzing data from different biological modalities. This benchmark aims to define critical issues in multi-omics research and guide future directions, ultimately promoting advancements in understanding biological processes through integrated and different omics data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。