arXiv:2604.06505cs.CLcs.AI2026-04被引 1

构建570万条生物医学结论生成数据集,评测大模型推理能力。

MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts

  • 基于结构化摘要构建证据到结论的推理任务
  • 570万条数据,支持跨领域子组分析
  • 揭示结论生成与摘要生成行为差异

大语言模型广泛用于复杂科研任务,但缺乏测试其从结构化生物医学证据中推导科学结论的能力资源。我们提出MedConclusion,一个包含570万条PubMed结构化摘要的大规模数据集,每条数据将摘要中的非结论部分与原始作者撰写的结论配对,提供自然的监督信号以支持证据到结论的推理研究。该数据集还包含期刊级别的元数据(如生物医学类别和SJR),可实现跨领域的子组分析。作为初步研究,我们在结论提示和摘要提示两种设置下评估多种LLM,并使用基于参考的指标和大模型作为裁判进行评分。结果表明,结论生成与摘要生成在行为上存在明显差异;当前自动指标下,强模型表现仍紧密聚集;裁判身份会显著影响绝对得分。MedConclusion为研究科学证据到结论的推理提供了可复用的数据资源。代码与数据已开源:https://github.com/Harvard-AI-and-Robotics-Lab/MedConclusion。

原文摘要 · Abstract (English)

Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific conclusions from structured biomedical evidence remain limited. We introduce $\textbf{MedConclusion}$, a large-scale dataset of $\textbf{5.7M}$ PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to-conclusion reasoning. MedConclusion also includes journal-level metadata such as biomedical category and SJR, enabling subgroup analysis across biomedical domains. As an initial study, we evaluate diverse LLMs under conclusion and summary prompting settings and score outputs with both reference-based metrics and LLM-as-a-judge. We find that conclusion writing is behaviorally distinct from summary writing, strong models remain closely clustered under current automatic metrics, and judge identity can substantially shift absolute scores. MedConclusion provides a reusable data resource for studying scientific evidence-to-conclusion reasoning. Our code and data are available at: https://github.com/Harvard-AI-and-Robotics-Lab/MedConclusion.

生物医学大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。