用大模型整合序列与结构信息,提升罕见基因变异分类准确率
Integrating Large Language Models for Genetic Variant Classification
- 融合多种大模型,结合基因与蛋白序列及结构特征进行分类
- 在ClinVar和ProteinGym数据集上表现超越现有工具,尤其擅长处理不确定变异
- 适合临床遗传诊断与精准医疗领域,助力个性化诊疗
遗传变异分类,尤其是意义未明变异(VUS),是临床遗传学与精准医学中的重大挑战。大型语言模型(LLMs)已成为该领域的变革性工具,能够发现传统方法难以察觉的复杂模式与预测线索,从而提升遗传变异致病性的预测准确性。本研究探究了先进LLMs——GPN-MSA、ESM1b和AlphaMissense的集成应用,这些模型利用DNA与蛋白质序列数据及结构信息,构建了全面的变异分类分析框架。我们采用标注完善的ProteinGym与ClinVar数据集评估这些集成模型,确立了新的分类性能基准。模型在具有挑战性的变异集上进行了严格测试,显著优于现有最先进工具,特别是在处理模糊和临床不确定性变异方面。研究结果表明,综合多种建模方法可显著提升遗传变异分类系统的准确性与可靠性。这些发现支持将先进计算模型应用于临床环境,可大幅提升遗传病诊断效率,推动个性化医疗向更精细、可操作的基因洞察迈进。
原文摘要 · Abstract (English)
The classification of genetic variants, particularly Variants of Uncertain Significance (VUS), poses a significant challenge in clinical genetics and precision medicine. Large Language Models (LLMs) have emerged as transformative tools in this realm. These models can uncover intricate patterns and predictive insights that traditional methods might miss, thus enhancing the predictive accuracy of genetic variant pathogenicity. This study investigates the integration of state-of-the-art LLMs, including GPN-MSA, ESM1b, and AlphaMissense, which leverage DNA and protein sequence data alongside structural insights to form a comprehensive analytical framework for variant classification. Our approach evaluates these integrated models using the well-annotated ProteinGym and ClinVar datasets, setting new benchmarks in classification performance. The models were rigorously tested on a set of challenging variants, demonstrating substantial improvements over existing state-of-the-art tools, especially in handling ambiguous and clinically uncertain variants. The results of this research underline the efficacy of combining multiple modeling approaches to significantly refine the accuracy and reliability of genetic variant classification systems. These findings support the deployment of these advanced computational models in clinical environments, where they can significantly enhance the diagnostic processes for genetic disorders, ultimately pushing the boundaries of personalized medicine by offering more detailed and actionable genetic insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。