用医学知识蒸馏构建可解释的药物推荐大模型
KEDRec-LM: A Knowledge-distilled Explainable Drug Recommendation Large Language Model
- 从医学知识库蒸馏知识,训练可解释药物推荐模型
- 构建首个可解释药物推荐数据集expRxRec,含临床试验与文献数据
- 适合医疗AI、药物研发人员研究可解释性推荐系统
药物发现是生物医学自然语言处理中的关键任务,但可解释的药物发现仍鲜有研究。大型语言模型在自然语言理解与生成方面表现出色,利用其进行可解释药物发现有望提升下游任务与实际应用效果。本研究整合开源药物知识图谱、临床试验数据及PubMed文献,构建了名为expRxRec的综合性可解释药物发现数据集。同时提出KEDRec-LM,一个经过指令微调的知识蒸馏型大语言模型,能从丰富医学语料中提取知识,实现药物推荐与推理生成。为推动该领域研究,本文将公开发布数据集与KEDRec-LM模型。
原文摘要 · Abstract (English)
Drug discovery is a critical task in biomedical natural language processing (NLP), yet explainable drug discovery remains underexplored. Meanwhile, large language models (LLMs) have shown remarkable abilities in natural language understanding and generation. Leveraging LLMs for explainable drug discovery has the potential to improve downstream tasks and real-world applications. In this study, we utilize open-source drug knowledge graphs, clinical trial data, and PubMed publications to construct a comprehensive dataset for the explainable drug discovery task, named \textbf{expRxRec}. Furthermore, we introduce \textbf{KEDRec-LM}, an instruction-tuned LLM which distills knowledge from rich medical knowledge corpus for drug recommendation and rationale generation. To encourage further research in this area, we will publicly release\footnote{A copy is attached with this submission} both the dataset and KEDRec-LM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。