将基因、转录和蛋白模型融合,提升分子属性预测效果
BioLangFusion: Multimodal Fusion of DNA, mRNA, and Protein Language Models
- 在密码子层面对齐三类序列模型嵌入,实现跨模态直接对应
- 五项任务中超越单模态基线,证明简单融合即可捕获多组学互补信息
- 无需额外训练,可无缝集成现有预训练模型,适合生物序列研究者
我们提出 BioLangFusion,一种将预训练的DNA、mRNA和蛋白语言模型统一融合为分子表示的简单方法。受分子生物学中心法则(从基因到转录本再到蛋白质的信息流)启发,我们在生物学意义明确的密码子层级(三个核苷酸编码一个氨基酸)对齐各模态嵌入,确保跨模态直接对应。BioLangFusion研究了三种标准融合技术:(i) 密码子级嵌入拼接,(ii) 受多实例学习启发的熵正则化注意力池化,(iii) 跨模态多头注意力——每种方法提供不同的归纳偏置以整合模态特异性信号。这些方法无需额外预训练或修改基础模型,可直接与现有基于序列的奠基模型集成。在五个分子属性预测任务中,BioLangFusion表现优于强单模态基线,表明即使简单的预训练模型融合也能以极低开销捕捉互补的多组学信息。
原文摘要 · Abstract (English)
We present BioLangFusion, a simple approach for integrating pre-trained DNA, mRNA, and protein language models into unified molecular representations. Motivated by the central dogma of molecular biology (information flow from gene to transcript to protein), we align per-modality embeddings at the biologically meaningful codon level (three nucleotides encoding one amino acid) to ensure direct cross-modal correspondence. BioLangFusion studies three standard fusion techniques: (i) codon-level embedding concatenation, (ii) entropy-regularized attention pooling inspired by multiple-instance learning, and (iii) cross-modal multi-head attention -- each technique providing a different inductive bias for combining modality-specific signals. These methods require no additional pre-training or modification of the base models, allowing straightforward integration with existing sequence-based foundation models. Across five molecular property prediction tasks, BioLangFusion outperforms strong unimodal baselines, showing that even simple fusion of pre-trained models can capture complementary multi-omic information with minimal overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。