用知识图谱增强大模型,让生物机制推理更准确可靠
Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- 结合知识图谱与大模型,构建多跳推理链来追踪生物机制
- 在多跳任务上超越现有方法,尤其在深层推理中表现突出
- 适合需要精准生物机制推理的研究者和医药研发人员
理解复杂的生物分子机制需要跨分子相互作用、信号级联和代谢通路的多步推理。尽管大语言模型(LLMs)在该领域展现潜力,但其应用受限于逻辑不一致和缺乏领域知识支撑。现有方法常加剧这些问题:推理步骤偏离生物学事实或无法捕捉长程机制依赖。为此,我们提出知识增强的长链思维(Knowledge-Augmented Long-CoT)框架,将LLMs与基于知识图谱的多跳推理链相结合。该框架通过引导式多跳遍历与剪枝构建机制链,并将其用于监督微调以提升事实一致性,再通过强化学习进一步优化推理可靠性与连贯性。此外,为克服现有基准在规模、范围和深层推理标注上的不足,我们引入PrimeKGQA,一个全面的生物分子问答基准。实验结果表明,虽然更大闭源模型在简单任务上仍表现良好,但在推理深度增加时,我们的方法展现出显著优势,在需结构化生物知识遍历的多跳任务上达到领先性能。这些发现凸显了结构化知识与先进推理策略结合在可靠、可解释生物分子推理中的有效性。
原文摘要 · Abstract (English)
Understanding complex biomolecular mechanisms requires multi-step reasoning across molecular interactions, signaling cascades, and metabolic pathways. While large language models(LLMs) show promise in such tasks, their application to biomolecular problems is hindered by logical inconsistencies and the lack of grounding in domain knowledge. Existing approaches often exacerbate these issues: reasoning steps may deviate from biological facts or fail to capture long mechanistic dependencies. To address these challenges, we propose a Knowledge-Augmented Long-CoT Reasoning framework that integrates LLMs with knowledge graph-based multi-hop reasoning chains. The framework constructs mechanistic chains via guided multi-hop traversal and pruning on the knowledge graph; these chains are then incorporated into supervised fine-tuning to improve factual grounding and further refined with reinforcement learning to enhance reasoning reliability and consistency. Furthermore, to overcome the shortcomings of existing benchmarks, which are often restricted in scale and scope and lack annotations for deep reasoning chains, we introduce PrimeKGQA, a comprehensive benchmark for biomolecular question answering. Experimental results on both PrimeKGQA and existing datasets demonstrate that although larger closed-source models still perform well on relatively simple tasks, our method demonstrates clear advantages as reasoning depth increases, achieving state-of-the-art performance on multi-hop tasks that demand traversal of structured biological knowledge. These findings highlight the effectiveness of combining structured knowledge with advanced reasoning strategies for reliable and interpretable biomolecular reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。