针对mRNA密码子层级结构,提出新型预训练方法提升模型性能与生成质量。
HELM: Hierarchical Encoding for mRNA Language Modeling
- 引入密码子层级编码机制,优化语言模型损失函数
- 在7个下游任务上平均提升8%,抗体区域标注任务表现更优
- 适合生物序列建模、mRNA设计与生成研究者使用
信使RNA(mRNA)在蛋白质合成中起关键作用,其密码子结构直接影响生物学特性。尽管语言模型在分析生物序列方面展现出潜力,但现有方法未能考虑mRNA密码子结构的层级性。本文提出一种名为HELM的新型预训练策略,将密码子层级结构融入语言模型训练中。HELM根据密码子同义性调节损失函数,使模型学习过程更符合mRNA序列的生物学现实。我们在多个mRNA数据集和任务上评估HELM,结果表明其在七个不同的下游属性预测任务以及抗体区域标注任务上,平均比标准语言模型预训练和现有基础模型基线提升约8%。此外,HELM还增强了语言模型的生成能力,能生成更贴近真实数据分布的多样化mRNA序列,优于非层级基线方法。
原文摘要 · Abstract (English)
Messenger RNA (mRNA) plays a crucial role in protein synthesis, with its codon structure directly impacting biological properties. While Language Models (LMs) have shown promise in analyzing biological sequences, existing approaches fail to account for the hierarchical nature of mRNA's codon structure. We introduce Hierarchical Encoding for mRNA Language Modeling (HELM), a novel pre-training strategy that incorporates codon-level hierarchical structure into language model training. HELM modulates the loss function based on codon synonymity, aligning the model's learning process with the biological reality of mRNA sequences. We evaluate HELM on diverse mRNA datasets and tasks, demonstrating that HELM outperforms standard language model pre-training as well as existing foundation model baselines on seven diverse downstream property prediction tasks and an antibody region annotation tasks on average by around 8%. Additionally, HELM enhances the generative capabilities of language model, producing diverse mRNA sequences that better align with the underlying true data distribution compared to non-hierarchical baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。