用语言模型对齐基因序列,提升远距离表达预测能力
Long-range gene expression prediction with token alignment of large language model
- 将基因序列与语言模型令牌对齐,实现符号化推理
- 在Geuvadis数据集上达0.65的斯皮尔曼相关系数,提升10%
- 可结合人类注释提示,支持上下文学习,适合基因调控研究
基因表达是影响人类表型变异和疾病的核心细胞过程。尽管深度学习模型有所进展,但近期基准测试显示其难以捕捉远距离调控规律。本文提出Genetic sequence Token Alignment(GTA),通过预训练大语言模型将基因序列特征与自然语言标记对齐,使语言模型以冻结状态进行基因序列的符号推理。该跨模态方法学习调控语法,并可引入基因特异性人类注释作为提示,实现现有模型无法支持的上下文学习。在淋巴母细胞系上训练,GTA在Geuvadis共轭细胞数据集上表现优于Enformer等前沿模型,斯皮尔曼相关系数达0.65,提升10%。此外,GTA能通过识别输入基因上下文中最具意义区域,增强对长程相互作用的解释性。GTA是一种利用预训练语言模型的新型跨模态基因表达预测方法,实现了从仅依赖序列数据的传统范式到融合语义理解的新范式转变。
原文摘要 · Abstract (English)
Gene expression is a cellular process that plays a fundamental role in human phenotypical variations and diseases. Despite advances of deep learning models for gene expression prediction, recent benchmarks have revealed their inability to learn distal regulatory grammar. Here, we address this challenge by leveraging a pretrained large language model to enhance gene expression prediction. We introduce Genetic sequence Token Alignment (GTA), which aligns genetic sequence features with natural language tokens, allowing for symbolic reasoning of genomic sequence features via the frozen language model. This cross-modal adaptation learns the regulatory grammar and allows us to further incorporate gene-specific human annotations as prompts, enabling in-context learning that is not possible with existing models. Trained on lymphoblastoid cells, GTA was evaluated on cells from the Geuvadis consortium and outperforms state-of-the-art models such as Enformer, achieving a Spearman correlation of 0.65, a 10\% improvement. Additionally, GTA offers improved interpretation of long-range interactions through the identification of the most meaningful sections of the input genetic context. GTA represents a powerful and novel cross-modal approach to gene expression prediction by utilizing a pretrained language model, in a paradigm shift from conventional gene expression models trained only on sequence data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。