arXiv:2409.09825cs.CLcs.AI2024-09被引 16

首个专用于基因-表型关联分析的大模型,提升医学遗传学研究效率。

GP-GPT: Large Language Model for Gene-Phenotype Mapping

  • 基于三百万条组学与医学遗传学术语,分两阶段微调训练
  • 在基因信息检索与关系判断任务中优于Llama2、GPT-4等主流模型
  • 适合基因组学、医学遗传学研究人员快速获取精准知识

预训练大语言模型在自然语言处理中表现卓越,但在生物医学领域面临多源组学数据复杂性与异质性的挑战。为此,我们提出GP-GPT,首个专注于基因-表型知识表征与组学关系分析的专用大模型。该模型在包含超过300万条基因组学、蛋白质组学及医学遗传学术语的综合语料库上,通过多个大规模验证数据集与科研论文进行两阶段微调。实验表明,GP-GPT在医学遗传信息检索、基因组信息获取及关系判定等任务中表现优异,显著超越Llama2、Llama3和GPT-4等先进模型。研究还揭示了生物因子实体在GP-GPT中的表征细微变化,为大模型推动基因-表型研究提供了新可能。

原文摘要 · Abstract (English)

Pre-trained large language models(LLMs) have attracted increasing attention in biomedical domains due to their success in natural language processing. However, the complex traits and heterogeneity of multi-sources genomics data pose significant challenges when adapting these models to the bioinformatics and biomedical field. To address these challenges, we present GP-GPT, the first specialized large language model for genetic-phenotype knowledge representation and genomics relation analysis. Our model is fine-tuned in two stages on a comprehensive corpus composed of over 3,000,000 terms in genomics, proteomics, and medical genetics, derived from multiple large-scale validated datasets and scientific publications. GP-GPT demonstrates proficiency in accurately retrieving medical genetics information and performing common genomics analysis tasks, such as genomics information retrieval and relationship determination. Comparative experiments across domain-specific tasks reveal that GP-GPT outperforms state-of-the-art LLMs, including Llama2, Llama3 and GPT-4. These results highlight GP-GPT's potential to enhance genetic disease relation research and facilitate accurate and efficient analysis in the fields of genomics and medical genetics. Our investigation demonstrated the subtle changes of bio-factor entities' representations in the GP-GPT, which suggested the opportunities for the application of LLMs to advancing gene-phenotype research.

基因表型大模型医学遗传

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。