arXiv:2606.01394cs.CL2026-06

用知识图谱增强大模型,从海量文献中自动挖掘可信的药物-疾病关系。

UniD$^3$: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning

  • 融合知识图谱与大模型生成,分两阶段提取并整合药物与疾病关系
  • 构建包含28,915条药物-疾病匹配等的六大知识图谱,关键任务F1达0.85以上
  • 生成结果经医生评审验证,支持可解释检索,适合医药研究与新药发现

系统刻画药物-疾病关系对新药研发和老药新用至关重要,但受限于生物医学文献的异质性和快速膨胀。现有数据集依赖人工标注且常不完整,纯大模型方法易产生幻觉且证据支撑弱。本文提出UniD$^3$,一个统一框架,结合大语言模型与知识图谱增强的检索增强生成(KG-RAG),实现药物-疾病匹配(DDM)、药物有效性评估(DEA)和药物靶点分析(DTA)的知识提取、组织与验证。UniD$^3$ 使用 Llama 3.3-70B 处理 157,849 篇 PubMed 文章,通过论文级抽取与图级合并的双阶段策略构建知识图谱。这些图谱支持基于 KG-RAG 的结构化数据生成,并通过外部基准测试、模糊匹配已有资源及临床医生评审进行验证。最终产出六套知识图谱与大规模数据集,包括 28,915 条 DDM、15,042 条 DEA 和超过 4,000 条 DTA QA 对。外部验证显示优异性能(DDM/DEA F1: 0.85–0.87;DTA F1: 0.82),临床评审确认高可靠性(AUROC = 0.90)。KG-RAG 增强模型优于独立大模型,且 UniD$^3$ 聊天机器人支持可解释、带引用的药物-疾病关系探索。该框架可扩展地将非结构化文献转化为高质量结构化知识,助力AI驱动的药物发现、再利用与精准医疗。

原文摘要 · Abstract (English)

Systematic characterization of drug-disease relationships is essential for drug discovery and repurposing, yet is hindered by the heterogeneity and rapid growth of biomedical literature. Existing datasets rely on labor-intensive curation and are often incomplete, while LLM-only approaches suffer from hallucination and weak evidence grounding. We introduce UniD$^3$, a unified framework that integrates Large Language Models with Knowledge Graph-enhanced Retrieval-Augmented Generation (KG-RAG) to extract, organize, and validate drug-disease knowledge across Drug-Disease Matching (DDM), Drug Effectiveness Assessment (DEA), and Drug-Target Analysis (DTA). UniD$^3$ processes 157,849 PubMed articles with Llama 3.3-70B and constructs knowledge graphs via a dual-stage strategy combining paper-level extraction with KG-level consolidation centered on drug and disease entities. These graphs support KG-RAG-based generation of structured datasets, evaluated through external benchmarks, fuzzy matching with curated resources, and clinician review. UniD$^3$ produces six knowledge graphs and large-scale datasets, including 28,915 DDM, 15,042 DEA, and over 4,000 DTA QA pairs. External validation shows strong performance (F1: 0.85-0.87 for DDM/DEA; 0.82 for DTA), with clinician review confirming high reliability (AUROC = 0.90). KG-RAG-augmented models outperform standalone LLMs, and the UniD$^3$ chatbot enables interpretable, citation-supported exploration of drug-disease relationships. UniD$^3$ provides a scalable, extensible framework for transforming unstructured biomedical literature into high-quality, structured drug-disease knowledge, supporting AI-driven discovery, repurposing, and precision medicine.

药物发现知识图谱RAG大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。