构建首个生物医学多跳多答案推理基准,揭示大模型在复杂关系推理中的短板。
BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain
- 基于PrimeKG构建多跳推理任务,涵盖一跳与两跳关系链
- 顶尖模型在两跳任务上仅达14.57%准确率,显示推理能力严重不足
- 适合关注生物医学AI推理、知识图谱与大模型评估的研究者
生物医学推理常需在药物、疾病、蛋白质等实体间跨越多重关联。尽管大语言模型日益普及,现有评测基准缺乏对生物医学领域多跳推理的评估能力,尤其无法处理一对一、多对多的关系。为此,我们提出BioHopR,一个基于全面的PrimeKG构建的新基准,用于评估结构化生物医学知识图谱中的多跳、多答案推理。该基准包含一跳和两跳推理任务,反映真实世界复杂性。对先进模型的评估显示,专有推理模型O3-mini在单跳任务上取得37.93%精度,在两跳任务上为14.57%,优于GPT4O及开源模型如HuatuoGPT-o1-70B和Llama-3.3-70B。然而所有模型在多跳任务中表现显著下降,凸显生物医学隐式推理步骤的挑战。BioHopR填补了该领域的评测空白,确立了新标准,并揭示专有与开源模型间的差距,推动生物医学大模型发展。
原文摘要 · Abstract (English)
Biomedical reasoning often requires traversing interconnected relationships across entities such as drugs, diseases, and proteins. Despite the increasing prominence of large language models (LLMs), existing benchmarks lack the ability to evaluate multi-hop reasoning in the biomedical domain, particularly for queries involving one-to-many and many-to-many relationships. This gap leaves the critical challenges of biomedical multi-hop reasoning underexplored. To address this, we introduce BioHopR, a novel benchmark designed to evaluate multi-hop, multi-answer reasoning in structured biomedical knowledge graphs. Built from the comprehensive PrimeKG, BioHopR includes 1-hop and 2-hop reasoning tasks that reflect real-world biomedical complexities. Evaluations of state-of-the-art models reveal that O3-mini, a proprietary reasoning-focused model, achieves 37.93% precision on 1-hop tasks and 14.57% on 2-hop tasks, outperforming proprietary models such as GPT4O and open-source biomedical models including HuatuoGPT-o1-70B and Llama-3.3-70B. However, all models exhibit significant declines in multi-hop performance, underscoring the challenges of resolving implicit reasoning steps in the biomedical domain. By addressing the lack of benchmarks for multi-hop reasoning in biomedical domain, BioHopR sets a new standard for evaluating reasoning capabilities and highlights critical gaps between proprietary and open-source models while paving the way for future advancements in biomedical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。