arXiv:2502.06173cs.LGcs.AI2025-02被引 2

让大模型学会评估自身不确定性,提升蛋白质互作预测的可信度。

Uncertainty-Aware Adaptation of Large Language Models for Protein-Protein Interaction Analysis

  • 用贝叶斯LoRA和集成LoRA量化模型不确定性。
  • 在多种疾病场景下保持高精度,且输出结果可信赖。
  • 适合需高可靠性的生物医学研究与精准医疗领域。

蛋白质-蛋白质相互作用(PPI)的识别有助于揭示细胞机制,尤其在神经退行性疾病、代谢综合征和癌症等复杂条件下。大型语言模型(LLMs)通过自动挖掘海量生物医学文献,在预测蛋白质结构和相互作用方面展现出巨大潜力;然而其固有的不确定性仍是获得可重复研究成果的关键挑战,这对生物医学应用至关重要。本研究提出一种面向PPI分析的不确定性感知大模型适配方法,基于微调后的LLaMA-3与BioMedGPT模型,结合LoRA集成与贝叶斯LoRA实现不确定性量化(UQ),确保对蛋白质行为推断的置信度校准。该方法在多种疾病背景下均实现了具有竞争力的PPI识别性能,同时有效缓解了模型不确定性,显著提升了计算生物学研究的可信度与可重复性。这些发现表明,不确定性感知的LLM适配有望推动精准医学与生物医学研究的发展。

原文摘要 · Abstract (English)

Identification of protein-protein interactions (PPIs) helps derive cellular mechanistic understanding, particularly in the context of complex conditions such as neurodegenerative disorders, metabolic syndromes, and cancer. Large Language Models (LLMs) have demonstrated remarkable potential in predicting protein structures and interactions via automated mining of vast biomedical literature; yet their inherent uncertainty remains a key challenge for deriving reproducible findings, critical for biomedical applications. In this study, we present an uncertainty-aware adaptation of LLMs for PPI analysis, leveraging fine-tuned LLaMA-3 and BioMedGPT models. To enhance prediction reliability, we integrate LoRA ensembles and Bayesian LoRA models for uncertainty quantification (UQ), ensuring confidence-calibrated insights into protein behavior. Our approach achieves competitive performance in PPI identification across diverse disease contexts while addressing model uncertainty, thereby enhancing trustworthiness and reproducibility in computational biology. These findings underscore the potential of uncertainty-aware LLM adaptation for advancing precision medicine and biomedical research.

大模型蛋白质互作不确定性量化精准医疗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。