arXiv:2606.16211cs.CL2026-06被引 1

构建多源生物医学推理数据集与框架,提升复杂知识整合能力

Weaving Multi-Source Evidence for Biomedical Reasoning: The BioMedHop Benchmark and BioWeave Framework

论文配图:Weaving Multi-Source Evidence for Biomedical Reasoning: The BioMedHop Benchmark and BioWeave Framework
图 1 · 摘自论文原文
  • 设计跨知识图谱、文献与网页的多源证据融合推理方法
  • 在10,045个实例上实现比基线高10.5%的整体性能
  • 支持小模型逼近GPT-4水平,适合医学AI研究者使用

生物医学问答日益需要对相互关联实体进行推理,但现有基准主要聚焦于考试式知识、文献理解或短程多跳推理,对源感知的图推理与证据拓扑构建关注不足。为此,我们提出BioMedHop,一个基于多源图结构证据的生物医学推理评测基准,包含10,045个实例,覆盖共享邻居匹配、交集推理、路径推理与计数任务,支持选项、开放式及数值回答形式。为支持该基准,我们进一步提出BioWeave框架,能够检索知识图谱路径,从文献和网络资源中收集线索,构建统一证据图,并通过实体级证据验证答案。实验表明,BioWeave在整体性能上优于对比方法,较强基线ToG-2提升10.5%;且可适配不同LLM骨干,使Qwen3-4B等小型模型达到接近GPT-4-Turbo的推理水平。

原文摘要 · Abstract (English)

Biomedical question answering (QA) increasingly requires reasoning over interacting entities, where supporting evidence is scattered across biomedical knowledge graphs, literature documents, and web-accessible resources. However, existing biomedical QA benchmarks mainly focus on exam-style knowledge, literature comprehension, or short-range multi-hop inference, leaving source-conditioned graph reasoning and evidence topology construction underexplored. To fill this gap, we introduce BioMedHop, a multi-source graph-grounded benchmark for evaluating biomedical reasoning over structured evidence topologies. BioMedHop contains 10,045 instances across KG, document, web, and hybrid evidence settings, covering shared-neighbor matching, intersection reasoning, path-based reasoning, and counting, with option-based, open-ended, and numeric count renderings. To support this benchmark, we further propose BioWeave, a source-aware reasoning framework that retrieves biomedical KG paths, gathers supporting clues from documents and web sources, assembles them into a unified evidence graph, and verifies answers through entity-level evidence support. Comprehensive experiments show that BioWeave achieves the best overall performance among compared methods on BioMedHop, outperforming the strong hybrid baseline ToG-2 by 10.5% in the overall average. Moreover, BioWeave consistently improves different LLM backbones and enables smaller models, such as Qwen3-4B, to achieve reasoning performance comparable to GPT-4-Turbo.

生物医学推理多源证据知识图谱大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。