构建1000个跨文档医学问答对,评测大模型多跳推理能力。
Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering
- 设计需两跳信息整合的医学问答数据集
- 顶级系统在概念准确率上达89.3%,远超基线67.4%
- 检索增强生成技术对性能提升至关重要
多跳问答在生物医学领域仍是重大挑战,要求系统整合多个来源的信息以回答复杂问题。为此,BioCreative IX MedHopQA共享任务旨在评估大语言模型(LLMs)在多跳推理方面的能力。我们构建了一个包含1000个具有挑战性的问答对的数据集,涵盖疾病、基因和化学物质,尤其关注罕见病。每个问题需通过整合两个不同维基百科页面的信息来解答。共有来自13支团队的48个系统参与。评估采用表面字符串匹配与概念准确性(MedCPT得分)双重标准。结果显示,基线模型与优化系统间存在显著性能差距:最先进系统在MedCPT指标上达到89.30%的F1分数,在精确匹配(EM)上达87.30%,而零样本基线分别为67.40%和60.20%。核心发现为:检索增强生成(RAG)及相关检索策略是取得优异表现的关键。此外,概念级评估在答案表面形式不同时仍能准确判断正确性。MedHopQA数据集已公开,以推动该领域持续发展。挑战材料:https://www.ncbi.nlm.nih.gov/research/bionlp/medhopqa,基准测试:https://www.codabench.org/competitions/7609/
原文摘要 · Abstract (English)
Multi-hop question answering (QA) remains a significant challenge in the biomedical domain, requiring systems to integrate information across multiple sources to answer complex questions. To address this problem, the BioCreative IX MedHopQA shared task was designed to benchmark in multi-hop reasoning for large language models (LLMs). We developed a novel dataset of 1,000 challenging QA pairs spanning diseases, genes, and chemicals, with particular emphasis on rare diseases. Each question was constructed to require two-hop reasoning through the integration of information from two distinct Wikipedia pages. The challenge attracted 48 submissions from 13 teams. Systems were evaluated using both surface string comparison and conceptual accuracy (MedCPT score). The results showed a substantial performance gap between baseline LLMs and enhanced systems. The top-ranked submission achieved an 89.30% F1 score on the MedCPT metric and an 87.30% exact match (EM) score, compared with 67.40% and 60.20%, respectively, for the zero-shot baseline. A central finding of the challenge was that retrieval-augmented generation (RAG) and related retrieval-based strategies were critical for strong performance. In addition, concept-level evaluation improved answer assessment when correct responses differed in surface form. The MedHopQA dataset is publicly available to support continued progress in this important area. Challenge materials: https://www.ncbi.nlm.nih.gov/research/bionlp/medhopqa and benchmark https://www.codabench.org/competitions/7609/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。