构建自动化基因组模型评测平台,推动基因组人工智能普及应用。
OmniGenBench: Automating Large-scale in-silico Benchmarking for Genomic Foundation Models
- 开发GFMBench框架,自动整合百万级基因序列与百余任务
- 覆盖四大基准数据集,实现基因组模型统一评测
- 开源工具+公开排行榜,助力科研与工业界快速迭代
近年来,以大语言模型为代表的AI技术进步,激发了对基因组基础模型(GFMs)突破的期待。隐藏在生命演化初期多样基因组中的‘自然代码’,蕴含着深刻影响人类与生态系统的能力。近期如Evo等基因组基础模型的突破,吸引了大量投资与关注,解决了长期存在的难题,使基因组的计算机模拟研究转向自动化、可靠且高效的范式。然而,在这一技术飞跃时代,基因组基础模型研究面临两大挑战:缺乏专用评测工具和多样化基因组的开源软件支持,阻碍了其快速发展与广泛应用。为此,我们提出GFMBench框架,专注于基因组基础模型的评测。该框架标准化评测套件并自动化执行多类开放源代码的基因组基础模型评估,整合来自四个大规模基准的数百万条基因序列及百余项基因组任务,推动基因组模型在基因理解与合成等长期难题上的广泛应用。此外,GFMBench以开源形式发布,提供友好的用户界面与多样化教程,适用于AutoBench以及复杂任务如RNA设计与结构预测。为促进基因组建模进一步发展,我们推出了公开排行榜,展示基于AutoBench的性能结果。GFMBench标志着基因组基础模型评测标准化与应用普及的重要一步。
原文摘要 · Abstract (English)
The advancements in artificial intelligence in recent years, such as Large Language Models (LLMs), have fueled expectations for breakthroughs in genomic foundation models (GFMs). The code of nature, hidden in diverse genomes since the very beginning of life's evolution, holds immense potential for impacting humans and ecosystems through genome modeling. Recent breakthroughs in GFMs, such as Evo, have attracted significant investment and attention to genomic modeling, as they address long-standing challenges and transform in-silico genomic studies into automated, reliable, and efficient paradigms. In the context of this flourishing era of consecutive technological revolutions in genomics, GFM studies face two major challenges: the lack of GFM benchmarking tools and the absence of open-source software for diverse genomics. These challenges hinder the rapid evolution of GFMs and their wide application in tasks such as understanding and synthesizing genomes, problems that have persisted for decades. To address these challenges, we introduce GFMBench, a framework dedicated to GFM-oriented benchmarking. GFMBench standardizes benchmark suites and automates benchmarking for a wide range of open-source GFMs. It integrates millions of genomic sequences across hundreds of genomic tasks from four large-scale benchmarks, democratizing GFMs for a wide range of in-silico genomic applications. Additionally, GFMBench is released as open-source software, offering user-friendly interfaces and diverse tutorials, applicable for AutoBench and complex tasks like RNA design and structure prediction. To facilitate further advancements in genome modeling, we have launched a public leaderboard showcasing the benchmark performance derived from AutoBench. GFMBench represents a step toward standardizing GFM benchmarking and democratizing GFM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。