arXiv:2504.18736cs.CL2025-04被引 4

构建医学论文证据检索基准,评估模型找支持证据的能力

EvidenceBench: A Benchmark for Extracting Evidence from Biomedical Papers

  • 用专家引导的流程生成假设并标注句子级证据
  • 模型表现远低于专家水平,差距明显
  • 提供10万篇标注数据,支持模型训练与研究

我们研究了在生物医学论文中自动寻找与假设相关的证据任务。这一过程对科研人员验证科学假设至关重要。为此,我们提出了EvidenceBench基准,其通过一种新流程构建:基于现有专家判断,自动生成假设,并对生物医学论文进行逐句标注以识别相关证据,确保完全忠实于人类专家意见。我们通过多组专家标注验证了该流程的有效性和准确性。我们在基准上评估了多种语言模型和检索系统,发现当前模型性能仍显著落后于专家水平。为展示流程可扩展性,我们构建了更大规模的EvidenceBench-100k,包含107,461篇带假设的完整标注论文,以支持模型训练与开发。两个数据集均开源:https://github.com/EvidenceBench/EvidenceBench

原文摘要 · Abstract (English)

We study the task of automatically finding evidence relevant to hypotheses in biomedical papers. Finding relevant evidence is an important step when researchers investigate scientific hypotheses. We introduce EvidenceBench to measure models performance on this task, which is created by a novel pipeline that consists of hypothesis generation and sentence-by-sentence annotation of biomedical papers for relevant evidence, completely guided by and faithfully following existing human experts judgment. We demonstrate the pipeline's validity and accuracy with multiple sets of human-expert annotations. We evaluated a diverse set of language models and retrieval systems on the benchmark and found that model performances still fall significantly short of the expert level on this task. To show the scalability of our proposed pipeline, we create a larger EvidenceBench-100k with 107,461 fully annotated papers with hypotheses to facilitate model training and development. Both datasets are available at https://github.com/EvidenceBench/EvidenceBench

证据提取医学AI数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。