arXiv:2608.06253cs.LG2026-08

用专精代谢组学的LLM构建可预测的代谢物图谱,提升疾病预测能力

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

  • 通过持续预训练与检索增强,打造代谢组学专用大模型
  • 在冠脉搭桥术后应激性高血糖预测中达AUC 0.8616
  • 生成的代谢物图谱可解释,适合生物医学研究者使用

代谢组学知识分散于异构数据源,难以转化为可预测表征。我们开发了MetaboLLM,一种通过持续预训练、监督微调和结构化检索优化的代谢组学专用大语言模型,并结合MetaboLLM-GIN,将生成的生化描述转换为代谢物图谱,用于患者级预测。在四种骨干模型架构下,MetaboLLM在代谢组学知识、关系推理和描述生成任务上均优于对应基线模型与医学适配模型,并在外部公开基准上成功迁移。MetaboLLM-GIN在冠状动脉旁路移植术后应激性高血糖预测(AUC 0.8616)和绝经后激素方案分类(AUC 0.8123)中表现最优,超越传统模型及非适配、无检索配置的生成图谱。模型可解释性分析揭示了生物学意义发现。结果表明,领域专用语言模型能将异构生化知识组织为可预测且可解释的代谢物图谱。

原文摘要 · Abstract (English)

Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations. We developed MetaboLLM, a metabolomics-specialized large language model adapted through continual pretraining, supervised fine-tuning, and structured retrieval, together with MetaboLLM-GIN, which converts generated biochemical descriptions into metabolite graphs for patient-level prediction using a graph isomorphism network. Across four backbone families, MetaboLLM outperformed corresponding base and medically adapted models on metabolomics knowledge, relational, and description tasks, and transferred to an external public benchmark. MetaboLLM-GIN achieved the highest AUC for stress hyperglycemia prediction after coronary artery bypass grafting (0.8616) and postmenopausal hormone-regimen classification (0.8123), outperforming conventional models, alternative graph constructions, and graphs generated from unadapted or non-retrieval LLM configurations. Model interpretation further produced biologically meaningful findings in both applications. These results show that domain-specialized language models can organize heterogeneous biochemical knowledge into predictive and interpretable metabolite graph representations.

代谢组学大模型图神经网络可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。