构建526亿词元生物通用预训练语料库,提升模型对分子到通路的跨尺度理解。
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

- 将分子、蛋白、基因组等分散生物数据统一为可训练语料
- 在固定模型下使多任务评测得分翻倍,各领域均有提升
- 适合需要生物知识理解的AI研发者与生信研究者
生物大语言模型的发展亟需能赋予模型真实生物学理解能力的训练语料。然而现有资源如分子数据库、蛋白质库、基因组注释、单细胞图谱和通路数据库,格式异构且未整合成统一语料。我们提出TheBioCollection,一个52.6B词元规模的预训练语料库,将小分子、蛋白质、基因组序列、细胞及通路数据统一转化为可训练形式。除整合现有数据外,还通过工具计算补充生物属性,并引入当前语料覆盖不足的新指令任务。配套推出TheBioCollection-Eval评测集,覆盖分子、蛋白、基因组、细胞及跨域的识别、生成与预测任务。在固定Gravity-16B-A3B模型架构下,基于TheBioCollection训练使整体评测得分翻倍,各领域均获提升,同时保持通用语言能力基本不变。
原文摘要 · Abstract (English)
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。