构建疾病中心的多跳推理数据集,提升医学大模型评估可信度。
MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering
- 基于两篇维基文章信息融合,设计需跨源推理的开放问答。
- 1000个专家标注问题,答案采用自由文本,避免选择题陷阱。
- 通过语义同义词增强和防污染机制,适合医学推理研究者使用。
评估大语言模型在生物医学领域的表现需要能区分推理与模式匹配、且随模型能力提升仍具区分力的基准。现有生物医学问答基准存在局限:多项选择题易导致模型通过排除法作答而非推理;广泛传播的考试类数据集则面临性能饱和与训练数据污染风险。多跳推理——即整合多个信息源得出答案的能力——是诊断支持、文献发现等临床任务的核心,但当前基准中仍不足。MedHopQA 是一个以疾病为中心的多跳推理基准,包含1000个专家精心标注的问答对,作为 BioCreative IX 共享任务发布。每个问题需综合两篇不同维基文章的信息,答案为开放式自由文本。金标准标注通过 MONDO、NCBI Gene 和 NCBI Taxonomy 的本体同义词集进行语义层面扩展,支持词汇与概念级评估。该数据集通过人工标注、筛选、迭代验证及大模型作为评判者的流程构建。为降低排行榜操纵和污染风险,1000个评分问题嵌入于公开可下载的10000个问题集(答案未提供)中,置于 CodaBench 平台。MedHopQA 不仅提供基准,还提供可复用的框架,未来生物医学问答数据集可优先考虑组合推理、抗饱和性与抗污染性三大设计原则。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabilities improve. Existing biomedical question answering (QA) benchmarks are limited in this respect. Multiple-choice formats can allow models to succeed through answer elimination rather than inference, while widely circulated exam-style datasets are increasingly vulnerable to performance saturation and training data contamination. Multi-hop reasoning, defined as the ability to integrate information across multiple sources to derive an answer, is central to clinically meaningful tasks such as diagnostic support, literature-based discovery, and hypothesis generation, yet remains underrepresented in current biomedical QA benchmarks. MedHopQA is a disease-centered multi-hop reasoning benchmark consisting of 1,000 expert-curated question-answer pairs introduced as a shared task at BioCreative IX. Each question requires synthesis of information across two distinct Wikipedia articles, and answers are provided in an open-ended free-text format. Gold annotations are augmented with ontology-grounded synonym sets from MONDO, NCBI Gene, and NCBI Taxonomy to support both lexical and concept-level evaluation. MedHopQA was constructed through a structured process combining human annotation, triage, iterative verification, and LLM-as-a-judge validation. To reduce leaderboard gaming and contamination risk, the 1,000 scored questions are embedded within a publicly downloadable set of 10,000 questions, with answers withheld, on a CodaBench leaderboard. MedHopQA provides both a benchmark and a reusable framework for constructing future biomedical QA datasets that prioritize compositional reasoning, saturation resistance, and contamination resistance as core design constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。