首个统一处理序列、结构与语言的生物通用模型,支持分子与蛋白跨模态理解生成。
BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language

- 用统一分词方案将序列、结构、语言映射到同一离散空间,单解码器原生支持多模态输入输出。
- 在80项任务中77项达领先或竞争水平,涵盖单/多实体、跨模态理解与生成任务。
- 适合生物信息学、药物设计、跨模态生成等需要整合多源数据的研究者使用。
我们提出BioMatrix,首个原生融合序列、结构与自然语言的多模态基础模型,覆盖分子与蛋白质。现有模型或局限于单一实体类型,或依赖适配器架构无法原生生成所读模态。BioMatrix通过统一分词方案,将分子序列(支持SMILES和SELFIES)、分子结构、蛋白序列、蛋白结构及自然语言映射至共享离散标记空间,实现所有模态在单一解码器下的统一读写,无需外部编码器、投影适配器或模态特定输出头。基于Qwen3语言模型(1.7B与4B参数),在3044亿个令牌上持续预训练,涵盖通用与领域文本、分子与蛋白的序列与结构视图,以及交织生物分子与科学文本的跨模态语料,并通过分子-蛋白与蛋白-蛋白互作数据链接不同实体。微调后在涵盖6类共80项任务的全面下游测试中,77项达到或超越当前最优性能,证明单一原生多模态通用模型可在广泛生物任务中媲美甚至超越专用方法。
原文摘要 · Abstract (English)
We present BioMatrix, the first multimodal foundation model that natively integrates sequences, structures, and natural language for both molecules and proteins within a single decoder-only architecture. Existing biological foundation models pursue native multimodality and broad entity coverage separately: those that fuse multiple modalities under a shared objective remain confined to a single entity type, while those spanning multiple entity types either omit explicit structural modeling or rely on adapter-based designs in which the model cannot natively generate the very modalities it can read. BioMatrix closes this gap by mapping molecular sequences (supporting both SMILES and SELFIES notations), molecular structures, protein sequences, protein structures, and natural language into a shared discrete token space through a unified tokenization scheme, so that all modalities are consumed and produced uniformly under a single next-token prediction objective -- without external encoders, projection adapters, or modality-specific output heads. Built upon the Qwen3 language model (1.7B and 4B), BioMatrix is continually pretrained on 304.4 billion tokens spanning general and domain-specific text, sequence and structure views of molecules and proteins, and cross-modal corpora that interleave biomolecular entities with scientific text and link distinct entities through molecule-protein and protein-protein interaction data. After tuning on a comprehensive suite of downstream applications covering 80 tasks across 6 categories -- encompassing single-entity and multi-entity understanding and generation tasks across and within modalities -- BioMatrix achieves state-of-the-art or competitive performance on 77 out of 80 tasks, demonstrating that a single, natively multimodal generalist model can effectively match or surpass specialized approaches across a wide range of biological tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。