arXiv:2509.24655cs.LGq-bio.GN2025-09

用双曲几何提升mRNA语言模型对生物层级结构的建模能力

HyperHELM: Hyperbolic Hierarchy Encoding for mRNA Language Modeling

  • 在欧氏主干上叠加双曲层,让mRNA表示贴近氨基酸层级关系
  • 多物种数据集上9/10任务优于欧氏基线,平均提升10%
  • 擅长长序列和低GC含量序列的分布外泛化,抗体区域标注准确率高3%

语言模型在蛋白质和mRNA等生物序列中的应用日益广泛,但其默认的欧氏几何可能无法匹配生物数据固有的层级结构。虽然双曲几何更适合表征层级数据,却尚未被引入mRNA语言建模。本文提出HyperHELM框架,在双曲空间中实现mRNA序列的掩码语言模型预训练。采用欧氏主干+双曲层的混合设计,使学习到的表示与mRNA-氨基酸间的生物学层级关系对齐。在多个跨物种数据集上,其在10项属性预测任务中有9项超越欧氏基线,平均性能提升10%,且在长序列和低GC含量序列的分布外泛化表现优异;在抗体区域标注任务中,相比层次感知的欧氏模型,准确率高出3%。结果表明,双曲几何是mRNA序列层次化语言建模的有效归纳偏置。

原文摘要 · Abstract (English)

Language models are increasingly applied to biological sequences like proteins and mRNA, yet their default Euclidean geometry may mismatch the hierarchical structures inherent to biological data. While hyperbolic geometry provides a better alternative for accommodating hierarchical data, it has yet to find a way into language modeling for mRNA sequences. In this work, we introduce HyperHELM, a framework that implements masked language model pre-training in hyperbolic space for mRNA sequences. Using a hybrid design with hyperbolic layers atop Euclidean backbone, HyperHELM aligns learned representations with the biological hierarchy defined by the relationship between mRNA and amino acids. Across multiple multi-species datasets, it outperforms Euclidean baselines on 9 out of 10 tasks involving property prediction, with 10% improvement on average, and excels in out-of-distribution generalization to long and low-GC content sequences; for antibody region annotation, it surpasses hierarchy-aware Euclidean models by 3% in annotation accuracy. Our results highlight hyperbolic geometry as an effective inductive bias for hierarchical language modeling of mRNA sequences.

mRNA建模双曲几何层次结构语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。