arXiv:2603.22327cs.IRcs.AI2026-03被引 1

用真实医学综述测试大模型,发现其能力不均衡且不可靠。

Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews

  • 构建自动化流程与专家标注数据集,覆盖1.6万篇论文,分阶段评估大模型在流行病学综述中的表现。
  • 模型在结构化数据提取上表现最差,平均字段级F1不足0.67,部分模型成本相差96倍。
  • 结果揭示大模型在关键任务中存在能力短板,不适合直接用于影响公共政策的医学决策。

系统性文献综述(SLRs)是高要求、高风险的科学知识整合形式,但目前仍缺乏对大型语言模型(LLMs)的有效评估框架。我们提出AgentSLR,一个大规模评估工具包,包含自动化综述工作流和涵盖16,248篇论文的专家标注数据集,用于评估大模型在流行病学系统性综述各阶段的表现。参考标注源自世界卫生组织优先病原体领域的同行评审研究,由领域专家生成。该工具包将每个综述阶段独立评估,并配备专用指标以实现针对性失败分析。我们测试了五种前沿推理模型,发现无一模型在所有任务中占优,表现出子任务专长,而整体基准掩盖了这种差异。结构化数据提取成为主要瓶颈,无模型字段级平均F1超过0.67。不同模型的估算成本差异高达96倍。记录的失败模式表明,当前模型尚不足以在流行病学中进行无监督部署,因结果可能影响公共政策。

原文摘要 · Abstract (English)

Systematic literature reviews (SLRs) are a demanding and high-stakes form of scientific knowledge synthesis that remains underspecified as an evaluation setting for large language models (LLMs). We introduce AgentSLR, a large-scale evaluation harness comprising an SLR automation workflow and an expert annotated dataset covering 16,248 articles, designed to test LLM capabilities across the stages of SLRs in epidemiology. Reference annotations were derived from peer-reviewed studies on WHO priority pathogens and produced by domain experts. The harness evaluates each review stage as a separate unit with dedicated metrics enabling targeted failure analysis. We evaluated five frontier reasoning models and found that no single model dominated across all tasks, showing sub-task specialisation often hidden by aggregate benchmarks. Structured data extraction is a major bottleneck, with no model exceeding an average field-level F1 of 0.67. Estimated costs vary substantially, by up to 96 times across evaluated models. Documented failure modes suggest that the evaluated models are not yet reliable enough for unsupervised deployment in epidemiology, where findings can inform public policy.

大模型评估医学综述流行病学知识提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。