首个评估大模型流行病学推理能力的基准,检验其从文献中推断疾病负担与干预效果的能力。
EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
- 构建三类渐进式任务,测试事实回忆、多步推理与不完整信息下的结论重构
- 15个模型表现普遍有限,多步推理是最大难点,规模不决定成败
- 支持细粒度诊断,适合研究医学AI证据推理与模型可解释性
可靠的流行病学推理需要整合研究证据以推断疾病负担、传播动态及干预效果。现有医疗问答基准多聚焦临床知识或个体层面推理,缺乏对基于证据的流行病学推断的系统评估。我们提出EpiQAL,据我们所知首个针对研究文献的流行病学问答诊断基准,包含三个子集,均来自涵盖多种疾病的开放获取文章。三个子集逐步测试事实回忆、多步推理和不完整信息下的结论重构,并通过质量控制流程(分类引导、多模型验证、难度筛选)构建。在15个涵盖开源与专有系统的模型上实验显示,当前大模型在流行病学推理上表现有限,多步推理挑战最大。模型排名随子集变化,规模不能预测成功。思维链提示对多步推理有帮助,但其他任务效果不一。EpiQAL提供关于证据锚定、推理与结论重构的细粒度诊断信号。
原文摘要 · Abstract (English)
Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clinical knowledge or patient-level reasoning, yet few systematically evaluate evidence-grounded epidemiological inference. We present EpiQAL, to our knowledge the first diagnostic benchmark for epidemiological question answering over research literature, comprising three subsets built from open-access articles across diverse diseases. The three subsets progressively test factual recall, multi-step inference, and conclusion reconstruction under incomplete information, and are constructed through a quality-controlled pipeline combining taxonomy guidance, multi-model verification, and difficulty screening. Experiments on fifteen models spanning open-source and proprietary systems reveal that current LLMs show limited performance on epidemiological reasoning, with multi-step inference posing the greatest challenge. Model rankings shift across subsets, and scale alone does not predict success. Chain-of-Thought prompting benefits multi-step inference but yields mixed results elsewhere. EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。