arXiv:2501.18794q-bio.GNcs.AI2025-01综述被引 11

用大模型提升罕见病致病基因发现,解决诊断难题。

Survey and Improvement Strategies for Gene Prioritization with Large Language Models

  • 采用多智能体与表型分类策略,分步优化基因排序。
  • 基础模型准确率近30%,分治策略显著提升效果。
  • 适合罕见病研究者、临床医生用于未解病例再分析。

罕见病诊断因患者数据有限和遗传多样性而困难,尽管变异优先排序技术有所进展,仍有许多病例无法确诊。尽管大语言模型(LLMs)在医学考试中表现优异,其在罕见遗传病诊断中的有效性尚未评估。我们针对基因优先排序任务对多种LLMs进行基准测试,结合多智能体系统与人类表型本体(HPO)分类,按表型特征和可解性对患者进行分组。随着基因集合规模增大,LLM性能下降,因此采用分治策略将任务拆分为更小子集。基线情况下,GPT-4表现最佳,正确排序致病基因的准确率接近30%。多智能体与HPO方法有助于区分易解与难解病例,凸显已知基因-表型关联及表型特异性的重要性。特定表型或明确关联的病例更易准确识别。然而,模型存在对研究充分基因的偏好及输入顺序敏感等偏差,影响优先排序效果。通过分治策略有效缓解了这些偏差。结合HPO分类、新型多智能体技术与优化的LLM策略,相较基线评估,显著提升了致病基因识别准确率。该方法可简化罕见病诊断流程,支持未解病例的重分析,加速基因发现,助力靶向诊断与治疗开发。

原文摘要 · Abstract (English)

Rare diseases are challenging to diagnose due to limited patient data and genetic diversity. Despite advances in variant prioritization, many cases remain undiagnosed. While large language models (LLMs) have performed well in medical exams, their effectiveness in diagnosing rare genetic diseases has not been assessed. To identify causal genes, we benchmarked various LLMs for gene prioritization. Using multi-agent and Human Phenotype Ontology (HPO) classification, we categorized patients based on phenotypes and solvability levels. As gene set size increased, LLM performance deteriorated, so we used a divide-and-conquer strategy to break the task into smaller subsets. At baseline, GPT-4 outperformed other LLMs, achieving near 30% accuracy in ranking causal genes correctly. The multi-agent and HPO approaches helped distinguish confidently solved cases from challenging ones, highlighting the importance of known gene-phenotype associations and phenotype specificity. We found that cases with specific phenotypes or clear associations were more accurately solved. However, we observed biases toward well-studied genes and input order sensitivity, which hindered gene prioritization. Our divide-and-conquer strategy improved accuracy by overcoming these biases. By utilizing HPO classification, novel multi-agent techniques, and our LLM strategy, we improved causal gene identification accuracy compared to our baseline evaluation. This approach streamlines rare disease diagnosis, facilitates reanalysis of unsolved cases, and accelerates gene discovery, supporting the development of targeted diagnostics and therapies.

基因优先排序大模型罕见病表型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。