arXiv:2503.23029cs.CL2025-03被引 19

用深度思考大模型构建生物医学知识图谱并提升跨文档推理能力

A Retrieval-Augmented Knowledge Mining Method with Deep Thinking LLMs for Biomedical Research and Clinical Support

  • 通过大模型自动构建生物医学知识图谱,整合海量文献信息
  • 跨文档问答任务中检索准确率提升20%,答案生成准确率提高25%
  • 适合临床医生制定个性化用药方案与科研人员发现研究空白

知识图谱与大语言模型(LLMs)是生物医学知识整合与推理的关键工具,有助于结构化组织科学文献并发现复杂语义关系。然而,现有方法面临知识图谱构建受限于专业术语、数据异构性及知识快速演化等问题,而大模型在检索与推理方面存在不足,难以揭示跨文档关联与推理路径。为此,我们提出一个流程:利用大模型从大规模文献中构建生物医学知识图谱(BioStrataKG),并建立跨文档问答数据集(BioCDQA)以评估潜在知识检索与多跳推理能力。随后引入集成渐进式检索增强推理(IP-RAR)框架,通过基于推理的集成检索最大化信息召回,并通过渐进式生成精炼知识,结合自我反思实现深度思考与精准上下文理解。实验表明,相较于现有方法,IP-RAR使文档检索F1分数提升20%,答案生成准确率提高25%。该框架帮助医生高效整合治疗证据以制定个性化用药方案,助力研究人员分析进展与发现研究空白,加速科学发现与决策进程。

原文摘要 · Abstract (English)

Knowledge graphs and large language models (LLMs) are key tools for biomedical knowledge integration and reasoning, facilitating structured organization of scientific articles and discovery of complex semantic relationships. However, current methods face challenges: knowledge graph construction is limited by complex terminology, data heterogeneity, and rapid knowledge evolution, while LLMs show limitations in retrieval and reasoning, making it difficult to uncover cross-document associations and reasoning pathways. To address these issues, we propose a pipeline that uses LLMs to construct a biomedical knowledge graph (BioStrataKG) from large-scale articles and builds a cross-document question-answering dataset (BioCDQA) to evaluate latent knowledge retrieval and multi-hop reasoning. We then introduce Integrated and Progressive Retrieval-Augmented Reasoning (IP-RAR) to enhance retrieval accuracy and knowledge reasoning. IP-RAR maximizes information recall through Integrated Reasoning-based Retrieval and refines knowledge via Progressive Reasoning-based Generation, using self-reflection to achieve deep thinking and precise contextual understanding. Experiments show that IP-RAR improves document retrieval F1 score by 20\% and answer generation accuracy by 25\% over existing methods. This framework helps doctors efficiently integrate treatment evidence for personalized medication plans and enables researchers to analyze advancements and research gaps, accelerating scientific discovery and decision-making.

知识图谱大模型临床支持多跳推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。