arXiv:2509.20702stat.APcs.AI2025-09被引 1

用大模型生成全基因组变异嵌入,助力精准医疗

Incorporating LLM Embeddings for Variation Across the Human Genome

  • 基于文本描述构建全基因组变异嵌入,覆盖89亿种可能变异
  • 在UKB数据集上验证嵌入可提升多基因风险评分性能
  • 开源资源支持大规模基因组研究,适合生物信息与医学研究者

大语言模型嵌入技术在生物数据中展现强大表征能力,但以往应用多局限于基因层面。本文提出首个系统性框架,为整个人类基因组的遗传变异生成嵌入表示。基于FAVOR、ClinVar和GWAS Catalog的注释,我们为89亿种可能的变异构建功能文本描述,并生成三类规模的嵌入:150万HapMap3/MEGA变异、9000万拟合的UK Biobank(UKB)变异,以及90亿种全部可能变异。采用OpenAI text-embedding-3-large及开源Qwen3-Embedding-0.6B模型生成嵌入。基准质量控制实验显示其对变异属性具有高预测准确性,验证了嵌入作为基因组变异结构化表示的有效性。进一步应用于真实世界的嵌入增强遗传风险预测,在UKB队列数据上展示出在多基因风险评分(PRS)中的优越性能。相关资源已公开于Hugging Face,为大规模基因组发现与精准医学提供基础支持。

原文摘要 · Abstract (English)

Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to date focus on gene-level information. We present one of the first systematic frameworks to generate genetic variant-level embeddings across the entire human genome. Using curated annotations from FAVOR, ClinVar, and the GWAS Catalog, we construct functional text descriptions for 8.9 billion possible variants and generated embeddings at three scales: 1.5 million HapMap3/MEGA variants, 90 million imputed UK Biobank (UKB) variants, and 9 billion all possible variants. Embeddings were produced using general purpose models including both OpenAI's text-embedding-3-large and the open-source Qwen3-Embedding-0.6B models. Baseline quality control experiments demonstrate high predictive accuracy for variant-level properties, validating the embeddings as structured representations of genomic variation. We further apply them to real-world embedding-augmented genetic risk predictions that demonstrate the performance of using LLM embeddings in polygenic risk score (PRS) style predictions over the UK Biobank cohort data. These resources, publicly available on Hugging Face, provide a foundation for advancing large-scale genomic discovery and precision medicine.

基因组大模型嵌入精准医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。